The book / contents
The whole book,section by section.
Read straight from the manuscript's own contents file, so it is what is actually written rather than what was planned. Every chapter and every section is listed below.
- 11 chapters
- 359 sections
- 135k words
- 37 figures
01 / jump to a chapter11 chapters · 359 sections
- 01The AI-Native Engineer19
- 02From LLMs to Agents51
- 03Context Engineering Fundamentals32
- 04Model Context Protocol50
- 05Spec-Driven Development24
- 06The SDD Workflow31
- 07SDD Frameworks Compared14
- 08Verification and Quality Gates66
- 09Agent Orchestration Patterns27
- 10Scaling AI-Native Engineering in Teams29
- 11AI-Native in Practice at Cogwheel16
02 / every sectionin full, nothing folded
00ForewordFront matter
01The AI-Native Engineer19 sections
- 01The Vibe-Coding Trap
- 02Two Engineers, One Task
- 03What AI-Native Engineering Is Not
- 04The Role Transformation: From Implementer to Orchestrator
- 05Why the Traditional Workflow No Longer Scales
- 06The New Center of Gravity: Intent, Constraints, and Verification
- 6.1Intent
- 6.2Constraints
- 6.3Verification
- 07Human-in-the-Loop: The Nonnegotiable Checkpoint
- 08What AI-Native Engineers Actually Do
- 8.1Before the work
- 8.2During the work
- 8.3Throughout
- 09The New Skill Stack: Eight Skills That Make You Irreplaceable
- 10The Cost of Not Adapting
- 11Career Implications: Why Software Engineering Is Becoming a Leadership Role
- 12Starting Out: How Junior Engineers Grow When AI Writes the Basics
- 13Summary
02From LLMs to Agents51 sections
- 01How LLMs Work: An Engineer's-Eye View
- 1.1Tokens
- 1.2The Transformer and attention
- 1.3In-weights knowledge versus in-context knowledge
- 1.4The context window is your working-memory budget
- 02Why LLMs Are Good at Code
- 2.1Training with verifiable outcomes: RLHF and RLVR
- 2.2Chain-of-thought and reasoning models
- 03Nondeterminism, Temperature, and Sampling
- 04The Model Landscape
- 4.1Open source versus closed source
- 4.2Small language models
- 4.3Multimodal models
- 4.4Embeddings and semantic search
- 05Where LLMs Excel and Where They Fail
- 06Benchmarks: What They Do and Don't Measure
- 07The API Layer: How Engineers Talk to LLMs
- 08What Makes an Agent: LLM, Tools, Context, Loop
- 8.1The LLM: The reasoning core
- 8.2Tools: How agents act on the world
- 8.3Context: Accumulated state
- 8.4The agentic loop: Reason, act, observe, adjust
- 8.5Beyond single agents
- 09Building an Agent in 50 Lines of Code
- 9.1Stopping conditions and iteration limits
- 9.2Agent frameworks
- 9.3Why this matters for users of coding agents
- 10The Agent Execution Model: Reason, Act, Observe
- 10.1The self-correction loop: Tests as observations
- 11Tool Use and Function Calling
- 11.1How function calling works under the hood
- 12Memory and State
- 12.1In-context memory
- 12.2Long-term memory through files
- 12.3Retrieved memory with RAG
- 12.4Agent-generated memory
- 12.5Context rot
- 13Agent Economics: Cost, Latency, and Scope
- 14Human-in-the-Loop
- 14.1HITL as a design principle
- 14.2Feedback as context enrichment
- 14.3The agent decides when it needs you
- 14.4Calibrating HITL to risk
- 15The AI Tool Landscape
- 15.1IDE-integrated coding assistants
- 15.2CLI and terminal-native agents
- 15.3Cloud and background agents
- 15.4Code review assistants
- 15.5Design and UI generation tools
- 15.6Choosing tools without chasing hype
- 16Summary
03Context Engineering Fundamentals32 sections
- 01From Prompt Engineering to Context Engineering
- 02The Core Problem: Right Context, Right Time, Limited Window
- 2.1The attention problem
- 2.2The lost-in-the-middle phenomenon
- 2.3Context rot
- 2.4The human working-memory parallel
- 2.5The right-time problem
- 2.6The two failure modes
- 03Mandatory and Optional Context Components
- 04A Taxonomy of Context Components
- 4.1System prompts
- 4.2User input and user-provided context
- 4.3Rules, style guides, and constraints
- 4.4Skills
- 4.5Tools
- 4.6Custom agents
- 4.7Environment context: runtime and metadata
- 4.8Conversation history and memory management
- 05The Context Assembly Process
- 06Autonomous Context Discovery: How Agents Explore Before They Write
- 6.1Code indexing and semantic search
- 6.2CLI-based exploration
- 6.3Language servers and symbolic code understanding
- 6.4What autonomous discovery does and doesn't replace
- 07The Agent Runtime Pipeline: From Request to Action
- 7.1Context assembly, planning, and tool execution
- 7.2The iterative loop: reason, act, observe, adjust
- 08The Golden Balance: Too Little Focus, Too Much Distraction
- 09Avoiding Context Rot
- 9.1Session management
- 9.2Token efficiency for large payloads
- 10Summary
04Model Context Protocol50 sections
- 01What MCP Is and Why It Matters
- 1.1What MCP actually does
- 1.2Why the MCP standard matters for engineers
- 02The MCP Architecture: Clients, Servers, and Transports
- 2.1Hosts, clients, and servers
- 2.2Transports
- 2.3Local servers versus remote servers
- 2.4How it connects in your daily workflow
- 03The MCP Ecosystem
- 3.1Configuring a server: a typical example
- 04The Economics of MCP
- 4.1What you pay for in an MCP session
- 4.2Controlling tool output size
- 05MCP Security Model
- 5.1Threats
- 5.2What Real Incidents Teach Us About MCP Security
- 5.3MCP Security Best Practices
- 5.4Four Security Principles
- 06Operating MCP in Production
- 6.1MCP Approved Registry
- 07Debugging MCP
- 7.1The server won't start
- 7.2The server connects but shows no tools
- 7.3The agent doesn't use the tool you expect
- 7.4A tool call fails with an error
- 7.5Reading tool-call logs
- 08MCP and Context Pollution
- 8.1Why Too Many MCP Servers Degrade Performance
- 8.2Agent-Specific MCP Server Configuration
- 09Building Your Own MCP Server
- 9.1When to build your own
- 9.2How hard is it?
- 9.3Writing good tool descriptions
- 10MCP in Agentic Pipelines
- 10.1What changes when there's no human in the loop
- 10.2Practical use cases for MCP in pipelines
- 10.3The reliability constraint
- 11Skills and CLI Scripts: When You Don't Need an MCP Server
- 11.1The pattern
- 11.2Skills that reference project scripts
- 11.3The CLI-first movement
- 11.4When to use skills and CLI tools versus MCP
- 12Advanced Tool Use: Scaling Beyond Static Tool Lists
- 12.1The three bottlenecks
- 12.2Tool search: Loading tool definitions on demand
- 12.3Programmatic tool calling: Code as orchestration
- 12.4Tool-use examples: Teaching by showing
- 12.5Matching the solution to the bottleneck
- 12.6Where the industry is heading
- 13Summary
05Spec-Driven Development24 sections
- 01The Root Problem: Why Do We Need SDD?
- 1.1Garbage In, Garbage Out: At Scale
- 1.2The Compounding Problem
- 1.3The Human Parallel
- 02You Already Write Specs
- 2.1When These Documents Are Missing
- 03The Case for Spec-Driven Development
- 3.1The Power of Learning by Writing a Spec
- 3.2Uncovering Unclear Requirements with a Spec
- 04The Economics of Writing Specs
- 05When SDD Is Not the Right Fit
- 06Specs as Durable Artifacts That Survive Tool Changes
- 6.1Specs and Organizational Knowledge
- 07SDD and Human-in-the-Loop
- 08Why Engineers Resist Writing Specs
- 09Plan Mode: The Bridge Between Vibe Coding and SDD
- 9.1Plan Mode in Practice
- 9.2Plan Mode and Spec-Driven Development
- 10Three Levels of SDD Maturity
- 10.11. Spec-First: Writing Specs Before Implementation
- 10.22. Spec-Anchored: Keeping Specs During Evolution
- 10.33. Spec-as-Source: The Spec as Primary Artifact
- 11What Good Specs Look Like
- 12Summary
06The SDD Workflow31 sections
- 01The Canonical Loop: Specify, Plan, Execute, Verify, Integrate, Learn
- 1.1Where Humans Must Stay in the Loop
- 1.2Stop Conditions and Escalation Rules
- 02Before the Loop: When You Need a Prototype First
- 2.1The Throwaway Branch Pattern
- 2.2Scratch Repo Prototypes
- 2.3A Concrete Example
- 2.4Consolidating Learnings into the Spec
- 2.5When to Skip the Prototype
- 03The Core Artifact Set
- 3.1spec.md: Intent, Constraints, Acceptance Criteria
- 3.2plan.md: Approach, Trade-Offs, Sequencing
- 3.3tasks.md: Atomic Tasks with Done Checks
- 3.4Optional Artifacts: risks.md, rollback.md, adr.md
- 04From Idea to Spec: Turning Intent into a Document
- 05From Spec to Plan: Turning Intent into Action
- 06From Plan to Tasks: Decomposition and Dependency Management
- 07Execution Patterns and Prompts
- 08Verification as Hard Rails: CI Gates That Make Output Shippable
- 09Integration and Deployment
- 10Learning and Iteration: Closing the Loop
- 11The Spec Lifecycle: What Happens After Implementation?
- 11.1The Two-Tier Model: What Works in Practice
- 11.2The Learning Phase as the Bridge
- 12SDD in Brownfield Projects
- 12.1The Core Challenge: Specs for Code That Was Never Specified
- 12.2Starting Small: The Spec Island Strategy
- 12.3Reverse-Engineering Intent with AI Assistance
- 12.4Handling Undocumented Behavior and Implicit Contracts
- 12.5Prioritizing What Gets Specified First
- 13Summary
07SDD Frameworks Compared14 sections
- 01Why Frameworks Exist: Structure, Consistency, Team Alignment
- 02GitHub Spec Kit
- 03OpenSpec
- 04BMAD Method
- 4.1Roles, Personas, and Guided Workflows
- 4.2BMAD in Practice
- 4.3BMAD in Brownfield Projects
- 4.4Custom Workflows, Custom Agents, and Adapting BMAD for Your Use Case
- 05Other Tools Worth Knowing
- 5.1Kiro
- 5.2Agent Skills
- 06Choosing the Right Framework
- 07Frequently Asked Questions on Using SDD in Production
- 08Summary
08Verification and Quality Gates66 sections
- 01The Bottleneck Moves
- 1.1Why Manual Review Doesn't Scale
- 1.2What This Chapter Promises
- 1.3The Two Big Questions
- 1.4Harness Validations and Self-Check Loops
- 02Before Merge, Layer 1: Deterministic Guardrails
- 2.1Linting and Formatting as a Baseline
- 2.2Dead Code and Unused Dependencies
- 2.3Type Checks and Compilation
- 2.4The Test Suite
- 2.5Mutation Testing as a Test Quality Gate
- 2.6Property-Based Testing for Invariants
- 2.7Static Application Security Testing (SAST)
- 2.8Dependency and Container Scanning
- 2.9Secrets Scanning
- 2.10Infrastructure as Code Scanning
- 2.11Supply-Chain Integrity
- 2.12Architecture Fitness Functions
- 2.13Performance and Bundle Budgets
- 2.14Query and Data-Access Performance
- 2.15Dependency Updates
- 2.16Contract Gates: API and Schema Breakage
- 2.17Accessibility and Internationalization
- 2.18Pull-Request Size and Scope Discipline
- 03Before Merge, Layer 2: LLM-Based Review
- 3.1Why Deterministic Checks Are Not Enough
- 3.2Adversarial Review Before the PR
- 3.3How an LLM Reviewer Works
- 3.4Two Ways to Organize the Review
- 3.5What to Look For
- 3.6Tuning the Reviewer to Your Codebase
- 04After Merge, Layer 3: Safe Deployment Strategies
- 4.1Feature Flags as the Default
- 4.2Blue-Green Deployments
- 4.3Canary Deployments
- 4.4Shadow Traffic and Dark Launches
- 4.5Progressive Delivery by User Segment
- 4.6Automated Rollback on Error Budget Burn
- 4.7Choosing Strategies
- 4.8Data Safety
- 05After Merge, Layer 4: Runtime Safety and AI-Powered Ops
- 5.1The Three Pillars, Briefly
- 5.2OpenTelemetry as the Common Substrate
- 5.3SLOs and Error Budgets
- 5.4Reducing Alert Fatigue
- 5.5Where AI Changes the Game
- 5.6AI Anomaly Detection on Metrics and Logs
- 5.7Post-Deploy Observability Analysis
- 5.8AI-Driven Root-Cause Analysis
- 5.9The Bug-Resolution Pipeline
- 5.10Closing the Loop Back to the Coding Agent
- 5.11Synthetic Monitoring and Chaos Engineering
- 06The Human in the Loop: Deciding What Needs You
- 6.1Two Mechanisms to Decide What Needs a Human
- 6.2The Deterministic Gate: Changes That Always Need a Human
- 6.3The AI Risk Scorer
- 6.4How Human Review Is Changing
- 6.5Reviewing the Reviewer
- 07Putting It Together
- 7.1A Reference Pipeline, End to End
- 7.2What to Adopt First
- 7.3Metrics for the System Itself
- 7.4Failure Modes
- 7.5The Trust Ladder
- 08Further Reading
- 09Summary
09Agent Orchestration Patterns27 sections
- 01Why Software Engineers Need to Understand Agent Orchestration Patterns
- 1.1Why Using a Single Agent Doesn't Scale
- 1.2The Tradeoffs of Orchestrating Multiple Agents
- 02One Agent, Bigger Jobs: Compaction and Scratchpads
- 2.1Compaction: The Automatic Fix That Gets Messy
- 2.2Scratchpads: Memory in a File You Control
- 03SDD as an Orchestration Pattern
- 3.1Using Different Models in Different SDD Phases
- 04The Generator-Critic Loop
- 4.1Worked Example: Adversarial Code Review
- 05The Goal Pattern: Looping Until the Goal Is Met
- 06The Ralph Loop
- 6.1How It Works
- 6.2What Makes It Work, and What Keeps It Safe
- 6.3Tools
- 07Running Multiple Agents in Parallel on the Same Codebase
- 7.1Context-Window Isolation
- 7.2Git Worktrees
- 7.3Isolated Cloud VMs
- 08The Subagent Pattern
- 8.1Ways to Arrange Work
- 8.2The Advisor: A Stronger Model for the Hard Parts
- 09Workflows: Orchestration as Code
- 9.1A Fixed Script: Letting the Agent Decide
- 10The Dark Factory: Where This Is All Going
- 10.1Two Modes of Working
- 11Final Considerations
10Scaling AI-Native Engineering in Teams29 sections
- 01Why Scaling to the Team Level Is a Different Problem
- 1.1The Bottleneck Moves, It Doesn't Disappear
- 1.2New Failure Modes That Only Show Up at Team Scale
- 1.3AI Is a Mirror and a Multiplier
- 02The AI-Native Team Canvas: Design How Your Team Adopts AI
- 2.1A Business Model Canvas, but for AI Adoption
- 2.2How to Run It
- 2.3The Eight Boxes
- 2.4Skill Gaps and System Gaps
- 2.5Usage Is Not Readiness
- 03Resistance Is Rational: Why Engineers Push Back
- 3.1The Three Objections You Will Hear
- 3.2Reframing the Craft
- 3.3What This Means for Leaders
- 04Standardize the Team Stack
- 4.1Standardize What Matters
- 4.2The Shared Context Is the Real Stack
- 4.3The Harness Gets Better Over Time
- 4.4Someone Has to Own It
- 4.5Leave Room to Experiment
- 05How an AI-Native Team Works
- 5.1Pair on the Agent, Not Just the Code
- 5.2Write and Review Specs as a Team
- 06The Team Gets Smaller and Tighter
- 6.1The Human Cost of Running a Fleet
- 6.2Agents That Work Overnight
- 6.3Product and Design Move Closer to the Code
- 07Measuring Adoption
- 08Summary
11AI-Native in Practice at Cogwheel16 sections
- 01Gridlock at Cogwheel
- 02The Mission: A Greenfield Service Inside a Brownfield World
- 03Week 1: Everything Except Signals
- 3.1A PRD, Written Together
- 3.2The AINE Canvas: How the Team Will Work
- 3.3Building the User Harness
- 3.4Wiring In Cogwheel's World
- 04Week 2: Splitting Into Stories
- 4.1Prototyping in Code, Not Pictures
- 4.2The Loop in Motion
- 05Weeks 3–4: Trusting the Code
- 5.1Watch One PR Go Through
- 06Weeks 5–8: A Fleet, Not a Hero
- 6.1Making It the Team's Way
- 6.2Measuring Honestly
- 07Week 9: Cogwheel Ships Pulse
03 / from here