Notes
Why My Automated Trading Strategies Haven’t Earned Real Money
What my trading engine's tests actually showed, why positive returns weren't enough, and the agentic paper experiment I'm building next.
Why I optimize LinkedIn once a year
What changed after I updated my LinkedIn profile, which fields I focus on, and why I leave it alone between annual reviews.
How I build 3D that still behaves like a web page
How junxiong.dev uses Three.js with native scroll, one active canvas, reduced motion and a readable fallback.
How agent integrations get built inside large organisations
Registries, auth brokering, generated tool schemas and the org failure modes that stall internal agent platforms — grounded in public specs.
Where your tokens actually go in a coding agent
How files and tool output fill a coding agent's context, what long-context studies show, and how I would keep a session focused.
How I think about hardware for local models
Bandwidth, capacity, KV cache and MoE active parameters — the four numbers that decide local inference hardware, with sources.
Prefill and decode want different machines
LLM inference has two phases with opposite bottlenecks. Which one you care about decides what hardware to buy.
Why the same prompt can cost more on the next request
How prefix caching actually bills, what silently breaks it, and a worked cost example across Claude Code, Codex, Cursor and the raw API.
What changes when you quantize a model
How quantization changes model size, speed and answers, and why weight memory and KV-cache memory need separate budgets.
/raw/<slug>.md. /llms.txt indexes them; /llms-full.txt is the whole site in one file.