Posts
Independent essays on AI-native work patterns, agent infrastructure, and what actually works when running agents as real tools.
-
The Five-Layer AI Stack Is an Inverted Pyramid
Jensen Huang's five-layer AI cake explains the dependency chain. An inverted pyramid reveals where participation expands and where market power concentrates.
-
Open Data Wasn't Missing. The Interface Was.
A 12-domain benchmark and Sentinel-2 experiment show where the Registry of Open Data MCP server works, where search fails, and what the interface unlocks.
-
Every Way to Attribute Cost on Amazon Bedrock, Live-Tested
A working map of every mechanism Amazon Bedrock has for attributing inference cost — resource tags, identity tags, and request metadata — with config examples, and where SaaS billing has to leave the AWS billing layer entirely.
-
Does Graphify Actually Help an AI Coding Agent?
A code graph earns trust only when it beats disciplined source search on a correct, source-backed task.
-
Why AI Coding Tools Are Getting Cheaper: Prompt Caching Explained
Updated:Prompt caching discounts repeated input, but the savings depend on exact prefix reuse, model-specific retention, and enough reads to repay each cache write.
-
Prompt Caching Is a Harness Design Problem, Not Just an API Toggle
A 186-request test shows why cache savings depend on stable tools, deterministic prefixes, deliberate breakpoints, routing keys, and measurement inside the agent harness.
-
Building AI Voice Agents with Amazon Nova Sonic and LiveKit: Patterns, Architecture, and a RAG Latency Surprise
Patterns that held up across a series of Nova 2 Sonic + LiveKit voice agents, a telephony architecture that took a real phone call, and a latency finding that contradicts the standard advice on voice RAG.
-
FLUX.1 LoRA fine-tuning on fal.ai is remarkably simple
A tested workflow for training a FLUX.1 style LoRA on fal.ai, applying it through the current API, and understanding where consistency still falls short.