How agent integrations get built inside large organisations
Registries, auth brokering, generated tool schemas and the org failure modes that stall internal agent platforms — grounded in public specs.
Where your tokens actually go in a coding agent
What really fills a coding agent's context window, why long sessions give worse answers, and what separates the engineers who get a speedup.
What to actually buy to run open models locally
Bandwidth, capacity, KV cache and MoE active parameters — the four numbers that decide local inference hardware, with sources.
Prefill and decode want different machines
LLM inference has two phases with opposite bottlenecks. Which one you care about decides what hardware to buy.
Prompt caching, and why the same prompt costs 10x more on some days
How prefix caching actually bills, what silently breaks it, and a worked cost example across Claude Code, Codex, Cursor and the raw API.
What quantization actually costs you
Formats, measured quality loss, the speed you get back, and why the KV cache is the part quantization never shrinks.
/raw/<slug>.md. /llms.txt indexes them; /llms-full.txt is the whole site in one file.