Why Raw Accessibility-Tree Counts Need Environment Controls
Two Windows hosts on one Chrome build returned different accessibility-tree node totals. The experiment shows why viewport and node roles must be reported.
Read story
Two Windows hosts on one Chrome build returned different accessibility-tree node totals. The experiment shows why viewport and node roles must be reported.
Read story
A measured Skynet field study of UIA, CDP, coordinate mapping, occlusion, and provider acceptance with corrected experiments and raw evidence.
Read story
I built RStackBench to test when browser agents look successful but fail the requested state: 288 conditions, 864 traces, and an open preprint.
Read story
Benchmark accuracy and operational reliability are different properties. Two 2026 papers show capability gains barely moved reliability.
Read story
Most machine-learning tutorials quietly assume NVIDIA. I wanted to know how far a consumer AMD Radeon RX 6600 could go on Windows without CUDA. I tested PyTorch through torch-directml, TensorFlow with the DirectML plugin, ONNX Runtime, matrix workloads, a full MLP training loop, and
Read story
An RX 6600 DirectML benchmark verified both LSTM nodes on DirectML, found CPU median latency lower, and withheld 1,062 unvalidated forecasts.
Read story
I wanted the small, direct feel of a raw Chrome DevTools connection in Firefox without pulling in a full automation framework. Firefox no longer exposes CDP, so I built a roughly 700-line Python adapter that gives my code a CDP-style Tab API over WebDriver
Read story
A proof-backed field note on separating deterministic risk gates from neural review, using a real Gemini 3.5 and ChatGPT research run and official 2026 guidance.
Read story
An AI agent can pass every eval and still take an action it was never allowed to take. Permission drift is a control-plane failure, not a model failure. Here is the fix.
Read story
A Skynet field note on why production AI agents need reliability contracts: idempotency, retries, traces, scoped tools, evals, and bounded blast radius.
Read story
Why a smarter LLM will not fix an AI system that fails at orchestration and verification: reliability comes from the weakest coordination layer, not the model.
Read story
A 2026 field note on why multi-agent AI fleets fail at the control plane, and how live probes, decoupled agents, fallback chains, and independent monitors keep capability alive.
Read story