The agent harness: build the loop, adopt the sandbox
The agent harness is four build-versus-adopt decisions wearing one name. Where owning the loop pays for itself, and where it quietly does not.
→ read_post()Long-ish posts about work we shipped — the decisions, the bugs, the numbers. No think-pieces, no listicles.
The agent harness is four build-versus-adopt decisions wearing one name. Where owning the loop pays for itself, and where it quietly does not.
→ read_post()Durable execution engines assume that re-running your code reproduces the same decisions. LLM agents break that assumption, and the recovery semantics you pick have real costs.
→ read_post()Raw agreement gives an LLM judge credit for guessing. Why Cohen's κ belongs on the dashboard instead, what it does not fix and how to run the position swap.
→ read_post()Conformal prediction gives a real coverage guarantee, but it assumes exchangeability and one shot per item. Retries, drift and calibration reuse quietly break it.
→ read_post()Character-level accuracy hides the failure that actually breaks document RAG. Why reading order deserves to be your headline parsing metric, and how to measure it without labels.
→ read_post()Streaming video-language models are graded on answer accuracy, but live deployments fail on response timing, false alarms and frame budget. What to measure instead.
→ read_post()Iceberg v3 makes row lineage mandatory, but the spec exempts rows updated via equality deletes — and that gap quietly corrupts any consumer treating _row_id as a stable key.
→ read_post()Third-party MCP servers put attacker-controlled text into your model context. What the disclosed incidents actually show, and the dull controls that work.
→ read_post()Hi, I'm the ImmovableTech assistant. Ask me about our services, past projects, or how to get in touch.