THE SCOREBOOK LIBRARY

Ideas worth
understanding.

What changed. Why it matters. How to use it. Source-linked analysis and practical engineering guides.

Latest updates

Read, understand, build

15 complete articles
Release analysis

Haiku 5.5 is 90% cheaper per token. Your bill needs a closer look.

Claude Haiku 5.5 pricing, tokenizer changes, vendor and independent benchmarks, and a practical plan to cut cost per accepted result.

Practical guide

Haiku 5.5 pricing: the 100K boundary, cache break-even, and routing math

Calculate Haiku 5.5 token costs, cache write amortization, long-prompt charges, and fallback costs with a tested offline calculator.

Practical guide

Migrate to Haiku 5.5: request changes, response handling, and a runnable example

Move from Haiku 4.5 to 5.5 with adaptive thinking, effort, token recounting, robust text parsing, and staging checks for tools and refusals.

Agent engineering

Agent Context Compaction, KV Cache, and Prompt Caching Explained

Understand three different memory mechanisms, calculate when compaction pays, and preserve agent state without mistaking cached tokens for durable memory.

Agent economics

Claude Agent Routing: Cost per Accepted Task Beats the Cheapest Token

Build a defensible agent routing policy with conditional recovery rates, verifier errors, effort sweeps, batch deadlines, and a tested offline calculator.

Deployment decisions

China vs USA LLMs: Choose the Deployment, Not the Flag

Compare Chinese and US LLM deployment options through model licensing, inference geography, data retention, modalities, and worked total-cost economics.

Model profile

Claude Haiku 5.5: a model profile for work you can verify

A practical Claude Haiku 5.5 model profile: specifications, effort settings, benchmark evidence, workload fit, limitations, and an adoption framework.

Release analysis

DeepSeek V4.1 Flash: Architecture, Pricing, and the Real Cost of Agents

A source-checked guide to DeepSeek V4.1 Flash, asymmetric inference, cache economics, benchmark limits, and comparisons with GPT, Claude, Gemini, and Qwen.

Build with AI

Build a document extraction pipeline that knows when to stop

An offline Python workflow for invoice extraction with strict JSON, field-specific source evidence, bounded repair, and a review queue—plus a plan for evaluating a live model adapter.

Model comparison

GPT-6.1 Sol vs Claude Sonnet 5.5: Choosing a Coding Agent Without a Fake Winner

A rigorous comparison of current coding evidence, token tariffs, reasoning budgets, and deployment evaluation for GPT-6.1 Sol and Claude Sonnet 5.5.

Product release analysis

GPT-6 Intelligent UI: What Changed and How to Evaluate It

A technical reading of ChatGPT's Intelligent UI release, with clear API boundaries, streaming tradeoffs, accessibility checks, and an evaluation plan.

Caching economics

GPT-6 Prompt Caching Economics: Engineer for Reuse, Measure Accepted Work

A rigorous guide to GPT-6.1 Sol prompt caching, Claude and Gemini cache economics, break-even calculations, prefix design, and tested cost accounting.

Workload economics

LLM Price Wars: Compare Workload Costs, Batch Deadlines, and Realtime Systems

Normalize AI workload bills across tokens, promotions, context thresholds, off-peak schedules, batch orchestration, and realtime audio before choosing a provider.

Multimodal engineering

Qwen3.8 Omni Flash: How to Deploy a Multimodal Agent Without Confusing the APIs

Qwen3.8 Omni Flash explained: native audiovisual reasoning, nonrealtime versus realtime APIs, verified pricing, context limits, and accepted-task economics.

Benchmark methodology

Terminal-Bench 4.0 and SWE-bench: How to Compare Coding Agents Honestly

Understand Terminal-Bench 4.0, SWE-bench, harness effects, paired statistics, and cost per resolved task—with tested Python and current source evidence.

THE INTELLIGENCE DESK

The evidence behind your next choice.

Coding, reasoning, agents and evaluation cost—with sources and configuration details.

Explore the leaderboard
Explore Scorebook
DiscoverLeaderboardCompare modelsLatest updatesThe journalPractical guidesMethodology