From practice
Notes
Short, evidence-backed pieces about things I built and measured – vision LLMs, agents, evals and what goes wrong along the way.
- 5 min read
Eight tools, zero write access: approval gates for AI agents as code
The useful question is not whether an agent gets it wrong, but what it can do when it does. How a tool contract, proposal-only writes and a version check limit that – and where the boundary stops helping.
- AI Agents
- Tool-Rechte
- Approval-Gates
- Audit Ledger
- 6 min read
LangGraph deep research: verify citations instead of just linking
How Abundance binds claims, admitted evidence and verbatim quotes in code – and why semantic verification deliberately measures in shadow mode only.
- LangGraph
- Deep Research
- Citations
- Claim Verification
- 6 min read
LLM evals with a small reference dataset: what 12 cases can do
What a versioned twelve-case eval can reliably protect, how three trials scoring 11/12, 10/12 and 11/12 are evaluated, and where the evidence stops.
- LLM Evals
- Golden Dataset
- Regression Testing
- Claim Verification
- 5 min read
Optical context compression: OCR-heavy PDFs as image bundles for vision LLMs
Why text as an image can be cheaper, how the pipeline with Mistral OCR and a sizing model is built, and what a 22-page test document shows: up to 85 % less context.
- Vision-LLM
- OCR
- MCP
- Token-Kosten