Posts
Independent essays on AI-native work patterns, agent infrastructure, and what actually works when running agents as real tools.
-
So What Is the Effective Cost for a Million Tokens on Frontier Models?
The sticker price on a model card is not what you pay. Measuring cache-hit ratios from months of real usage across Sonnet 5, Opus 5, and GPT-5.6 puts the effective cost per million tokens 70-95% below list price.
-
What Actually Differs Across Search APIs Built for Agents
Updated:Exa, Parallel, Perplexity, and AgentCore span raw retrieval, generated answers, managed research, and monitoring. Comparing them starts with separating those layers.
-
Search for Agents Is Becoming Its Own Market
Updated:Four search APIs built for AI agents don't agree on a winner — and their funding, product surfaces, and billing units define an infrastructure market, not a feature comparison.
-
A 200 Response Does Not Prove Search Worked
HTTP success proves transport. A credible agent-search evaluation separately tests the contract, evidence, tool calls, lifecycle cleanup, and cost.
-
Search Is a Pipeline, Not a Tool
Agent search works as a five-stage pipeline: discovery, extraction, verification, synthesis, and monitoring. Treating it as one tool hides the failures that matter.
-
Wiring Codex and ChatGPT Desktop to OpenAI Models on Amazon Bedrock Runtime
A tested shared configuration for running Codex CLI and the ChatGPT desktop app against global OpenAI model profiles on Amazon Bedrock Runtime.
-
Open Weights Catch the Last Frontier, Not the Moving One
Across 19 model snapshots, open-weight leaders reached an earlier proprietary frontier in 24–58 days on coding, intelligence, and agentic benchmarks—but remained behind the live frontier.
-
The Gateway Controls the Request. The Router Chooses the Model.
Model routing is growing because model choice has become a runtime decision. Gateways govern traffic; routers decide which model should answer.