Skip to main content

Chapter 6: Building with Coding Agents

How do we use coding agents to explore, build and maintain Cloudflare applications well?


A coding agent can make an architectural disagreement cheap enough to test. Give it a representative query, two candidate data models and a way to measure the deployed request path, and it can help build the experiment that would otherwise remain on the backlog. It can also produce an impressive application around the wrong data model. Both outcomes begin with code that runs.

The opportunity is to shorten the distance between a question and trustworthy evidence. More of the work that supports good engineering becomes affordable: reproductions, migration rehearsals, small comparison implementations, unfamiliar-code walkthroughs and tests of awkward failure paths. Feature implementation benefits too, but shipping more code is only one use of the capacity.

Cloudflare is an interesting platform for this approach. Workers, bindings and local runtime tools give an agent relatively compact things to inspect and exercise. Managed state and global deployment also hide important behaviour from a local success. A generated application can pass every test its author thought to write while choosing the wrong authority for a quota, sharing a production binding or losing accepted work after a disconnect.

The development process has to expose those mistakes. This chapter concerns using agents to build software, including applications with no AI feature. Chapter 19 covers agents inside the application itself.

Make the uncertain decision executable​

The example follows a question through an experiment to a decision. It previews services explained later: D1 stores relational records, R2 stores files, and a Durable Object combines one entity’s state with code that coordinates changes. Chapters 7 and 12–14 develop those services; here, compare the evidence each design produces and the work it leaves the team to own.

Consider a document intake service used by several organisations. Users upload files, see processing progress and approve a particular version for delivery to a recipient. R2 holds the files; metadata and acceptance records need a durable home. Later work extracts content and sends notifications. A document can be large, a user can retry, and an organisation must never read another organisation's records.

Before asking an agent to build the service, identify the decision most likely to invalidate it. Perhaps the team expects one D1 database per tenant but has never tested its largest customer's reporting workload. Perhaps it assumes uploads can pass through the application handler without excessive memory use. A polished upload screen resolves neither uncertainty.

Ask for a narrow experiment with an explicit outcome: run the representative report against realistic data, record its query plan and rows scanned, and measure the deployed path. Or compare streaming through a Worker with a direct R2 upload followed by server-side verification. Use the same payloads and acceptance rules for each candidate. The agent can build the comparison, generate synthetic data with the relevant distribution and explain the results; the team decides whether the measured behaviour meets the requirement.

Do not turn a comparison into two complete applications. Implement only the path that distinguishes the choices, and give the experiment a stopping condition. If both approaches meet the requirement, operational simplicity may settle the decision. If neither does, more frontend work will not help.

Follow the experiment to a decision​

Start with a user retrying a submission, with the same operation ID, because the first response never arrived. Both requests must resolve to one job. Reusing that operation ID for a different document must be rejected. The service also needs relational reporting across a tenant's jobs, but does not yet need live collaboration.

Ask the agent for two small candidates. D1 can enforce one job per tenant and operation ID with a unique constraint. A Durable Object can own each tenant-and-operation pair, keeping the accepted document identity in its private storage. The question is which design meets the requirement with less work to operate.

Set the stopping rule before implementation: both candidates must survive the same concurrent requests and interrupted handoffs, then meet the team's latency and capacity targets on the deployed path. Hold the fixtures constant, including a busy tenant and reuse of one operation ID for different documents. Neither candidate needs a user interface. Check the fixture generator against known totals so synthetic data does not quietly make the difficult tenant disappear.

Suppose the experiment produces the following observations. These are illustrative outcomes, not measurements of either Cloudflare service:

ObservationConsequence for this design
Both candidates keep one job per operation, reject conflicting reuse and meet the workload targetsEither can handle acceptance for this workload
D1 can serve the required report from its accepted-job recordsThe reporting path already has its authoritative data
The object candidate publishes accepted state into a reporting storeThat candidate adds a handoff whose lag and recovery need ownership

Choose D1 for this version. The object prototype has earned its cost by showing that its extra coordination capability does not address a current requirement. This conclusion would change if one document needed an active owner ordering live decisions among collaborators, or if the measured workload exposed a different constraint. Record that reconsideration trigger alongside the result.

Retain the common tests, deployment configuration and decision record; discard the losing implementation. The agent made a comparison affordable. The engineer decided which obligations the application should carry.

Use an existing system as a comparison target​

A migration offers another source of evidence: software that already serves users. Ask the agent to run the old and proposed paths against the same retained cases, then classify disagreements. For a Worker replacing a regional handler, compare response meaning, authorisation, durable writes and retry behaviour. A matching JSON response can hide a second notification or a missing record.

Normalise only differences the contract permits, such as a generated trace ID. Keep intended behaviour changes in a separate list so the agent does not “fix” the replacement by reproducing a known defect. Run both paths against isolated copies and test receivers; replaying production requests against two writers can perform the real operation twice. Include awkward records and failures, not just the requests easiest to capture.

A reference implementation can make a large problem divisible. Anthropic's compiler experiment used GCC as a comparison target to isolate failures when agents had converged on the same large task. That is a useful mechanism to adapt to migrations, without treating a compiler capability experiment as evidence that an application migration is safe. Equivalence establishes compatibility with the retained cases; it does not establish that the old behaviour was correct.

Use the passing comparison in a cutover rehearsal. Chapter 27's transfer of write ownership gives the agent a concrete sequence to exercise: stop the old writer, reconcile outstanding changes and check what rollback would do with writes accepted by the replacement. An agent can automate that rehearsal and report which operation IDs failed to transfer.

Give the agent a system it can understand​

A coding agent combines a model with tools for reading, changing and running software. The surrounding machinery is often called a harness: it supplies context, invokes tools, retains progress and decides when work continues or stops. A capable model with an unusable test environment spends its time guessing; a working environment gives its mistakes somewhere to become visible.

Start with the repository a new engineer would need. One command should run the relevant checks. Another should start the application with known test data. The architecture description should identify the state owners and the places where authority changes. An agent can help create this setup, but its first successful run must be inspected before it becomes the team's evidence source.

OpenAI's account of its internal coding-agent project describes making repository knowledge and application behaviour accessible to agents. It is a useful engineering case, with a greenfield product and substantial supporting investment; its estimated productivity gain should not become a forecast for an unrelated team.

Make understanding a deliverable​

An agent can also shorten the investigation before a change. Ask it to trace one request from authentication through state mutation to dispatch, with source locations and the calls and records observed in a fixture. Follow the references yourself at the points where authority or persistence changes. A source-linked walkthrough gives the next engineer a way to check the explanation without reconstructing the conversation.

For a difficult state machine, a small interactive explorer can let the team pause an operation, lose its response and inspect the remaining state. Simon Willison's practitioner examples use source walkthroughs and interactive explanations to make unfamiliar code easier to understand. Applied here, the important requirement is that the explorer's transitions agree with implementation or captured traces. An animation can make an incorrect model unusually persuasive.

Keep the walkthrough or explorer when another engineer can use it to predict a failure outcome or plan a change. Discard a polished explanation that cannot be checked. Understanding is a useful result even when the investigation concludes that the existing design should stay.

Keep durable context close to the code​

Use a short repository entry document, such as AGENTS.md, to point to the architecture, commands, conventions and constraints that govern the work. Put detailed reasoning beside the component or decision it concerns. Repeatedly pasting a long history into chat creates another version of the architecture that someone must keep consistent.

For a Worker project, the context should identify the compatibility date, package versions, configured bindings and intended execution mode. A recommendation from current documentation may require a newer runtime contract than the project uses. Keep the supported contract explicit; let a compatibility-date upgrade be a separate, tested change.

Cloudflare publishes documentation in agent-friendly formats, skills and MCP servers. These help an agent find platform information and, where granted, use operational tools. Select sources for the question at hand: a runtime API reference for its behaviour, a limits page for capacity, and an application decision record for why a particular service was chosen. A product tutorial does not know the organisation's acceptance rules.

A skill is reusable task guidance. It can explain how this team tests a Worker or reviews a migration, but the correctness of its advice still needs maintenance. When a failure recurs, improve the nearest durable control: a regression test for behaviour, a lint rule for a mechanical restriction, or an architecture note for a decision that requires judgement. Adding another paragraph to every prompt is often the least effective repair.

Specify the contract that matters​

A useful brief states the externally visible outcome, the invariants to preserve and the resources the agent may change. Name the decisions that require review, such as a new authority boundary or destructive migration. Leave room for the agent to propose a simpler implementation within those limits.

For the intake service, this is enough to begin a bounded slice:

DecisionContract for the slice
IdentityDerive tenant identity from verified credentials and authorise the requested resource within that tenant
AcceptanceAn accepted response means the required file has been verified and a durable processing obligation exists
RetriesA repeated submission with the same operation identifier and payload resolves to the original job; conflicting reuse is rejected
StorageR2 owns immutable file versions; D1 owns accepted jobs and dispatch obligations
ProcessingDispatch can fail without losing the accepted obligation; retries cannot silently create another logical job
ScopeImplement upload completion and status retrieval before extraction, approval or notification
EvidenceDemonstrate cross-tenant rejection, interrupted dispatch and retry recovery against observable state

Dispatch is the handoff from an accepted job to background processing. The brief deliberately exposes a cross-service boundary: an R2 object and a D1 transaction do not commit together. The design must choose an order and handle interruption. Uploading first can leave an unreferenced object; accepting first can leave a job whose bytes never arrived. The agent should surface this decision before filling it with code.

Challenge that promise before accepting the implementation. Reuse the original upload credential after success and check that the accepted bytes cannot change. Pause acceptance while abandoned-file removal runs, then resume it and inspect the result. An accepted job must retain its verified file. These tests expose requirements that a successful upload demonstration misses; Chapter 23 develops the storage lifecycle that enforces them.

Scale the brief to the decision. A pagination bug with a reproducible request and a fixed expected result is a useful bounded repair for an agent to finish and submit for review. Choosing what “accepted” promises across storage and dispatch needs a product decision first. A long specification is unnecessary for the former and cannot substitute for judgement in the latter.

Design the evidence before accepting the implementation​

An agent can write useful tests. It can also write tests that formalise its own misunderstanding. Start from the required behaviour and check observable results, such as returned data and stored records.

Suppose the implementation and its test both trust a tenant identifier in a request body. The test may prove perfect agreement between two wrong pieces of code. Start the security test from two authenticated tenants and an ownership rule, then observe what each can retrieve. The database contents and returned response establish the outcome.

For a retry, send the same operation both sequentially and concurrently and check that one job exists. Then reuse the identifier with a different final document version or requested action and check that the conflict is rejected, including when conflicting requests arrive together. Also lose the first response after commit and check that a retry recovers the accepted result.

For interrupted dispatch, force the queue call to fail after acceptance, restart the application path and verify that the durable obligation can still be found. The test should fail if the implementation forgets the pending job, even when the HTTP handler returned the expected status.

Where practical, demonstrate the test against the known fault before using it to approve the fix. A failing test establishes that the check reaches the defect; a passing test afterwards establishes that this version resolves that scenario. It is a stronger signal than asking the agent whether its test suite is comprehensive.

Use a failure to improve the design​

Imagine the first generated handler reads for an existing operation, finds none and inserts a job. Its sequential retry test passes. A reviewer asks what happens when two requests both complete the lookup before either inserts. This is a specific counterexample, and it changes the work: the agent must exercise concurrent requests and put uniqueness inside the database operation.

For the D1 candidate, a unique constraint prevents two rows for the same tenant and operation identifier. It does not settle what a conflicting reuse means. The handler must compare the retained request identity with the new request and return either the original job or a conflict. Record the obligation to dispatch the job alongside its acceptance. This pending-work record is an outbox; putting both records in one transaction means they succeed or fail together. Otherwise success can leave a job without its dispatch obligation.

D1's transactional batch() can commit those writes together; separate calls from a Worker do not form one transaction. Chapter 13 develops that boundary.

For the Durable Object candidate, route every retry for the same tenant and operation to the same object, and retain the accepted request in its storage. SQLite constraints there apply within that object, not across the namespace. The object's identity brings requests to one owner; the storage rules preserve the decision.

Chapter 7 explains the remaining concurrency boundary. An object handler that awaits an external service between checking and changing state can let another request interleave. Keep that wait in the concurrency test, so a move to Durable Objects cannot hide the original race.

A useful review returns an execution sequence the implementation did not consider. Turn that sequence into a regression check, and ask the agent to explain which authority now prevents it. More reviewers help only when they add another way to test the claim.

Test where the failure can happen​

Use the development modes from Chapter 5 deliberately. Pure logic can be checked on its own. For Worker code, Cloudflare's Workers Vitest integration exercises runtime APIs and local bindings. A mock that always returns the expected job cannot reveal a broken SQL constraint or incorrect binding call. Deployed tests establish the real dependency path, placement and resource headroom.

For the intake service, a local test can show that job creation and the outbox record share a transaction. For the object candidate, repeat the request after discarding the in-memory instance while retaining its storage: the accepted decision must survive reconstruction. Neither test proves acceptable upload behaviour under deployed concurrency or that every resource binding targets staging. Run those checks where the relevant constraint exists.

Give the agent structured test failures, a local browser and a way to inspect test state. On the deployed intake path, carry the operation ID through Worker logs, D1 records and queued work so an interrupted handoff can be reconstructed. Chapter 21 explains how Workers Logs and tracing contribute to that evidence. Give the agent scoped access to the relevant diagnostics; keep credentials and personal data out of fixtures and output.

A browser check has a different job from an API check. The API may have accepted the file correctly while the interface still reports failure and invites a duplicate upload. Have the agent exercise the actual user journey, including reload and reconnect, and inspect the result yourself when the distinction depends on usability or domain judgement.

Keep acceptance checks independent​

An agent that edits the application, test expectations and release configuration can make a failing change look successful by changing the question. Sometimes tests are obsolete and should change; that decision must be visible.

Review changes to acceptance fixtures, permission checks, test skips and CI rules separately from the feature they assess. Run required checks through a trusted pipeline against the proposed revision. A status message in the conversation is a report, and should link to the command, result and revision that support it.

A second agent can challenge the design, inspect an untouched requirement or try to construct a counterexample. Give it a different task: “Find a sequence that produces two accepted jobs” is more useful than “Check whether this looks good”. Different prompts or models can broaden the search, but their agreement is not proof of independent reasoning. A reproducible failure or a checked invariant carries more weight than several approving summaries.

Anthropic's work on long-running application development found value in separating generation and evaluation, while also noting that an evaluator can remain too generous. The practical lesson is to give reviewers criteria and observable outcomes, then calibrate their findings against cases whose answers are known.

Make autonomous work recoverable​

A task that spans several sessions needs an explicit handoff. Retain the goal, decisions made, current revision, checks completed, open failures and the next experiment. On resumption, inspect the repository and rerun the relevant observation rather than treating the handoff as proof that nothing changed.

For the intake service, “dispatch recovery is finished” is too weak. “Revision X passes the interrupted-dispatch test; deployed queue configuration remains untested” tells the next worker where evidence ends. Store such progress with the task, and remove temporary detail when it is replaced by a decision record or a regression test.

Measure the delay between a change and useful feedback before adding more agents. If fixture reset or deployment dominates each attempt, improving that path may buy more exploration than extra parallel runs. Fast selective checks support iteration; the complete required checks still govern acceptance.

Continuous iteration is useful when a failure has a concrete signal and the agent can act on it. A loop can fix a type error, exercise the browser and try again. It becomes wasteful when it repeatedly changes code without learning why the check fails. Set a budget for the experiment and a stopping condition: an unresolved requirement, inaccessible dependency, repeated failure without a new hypothesis, or a result needing domain judgement.

The right response to a stalled run may be better evidence, a smaller task or direct human work. Increasing the number of attempts does not supply a missing business rule.

Parallelise around contracts​

Multiple agents help when work can be evaluated independently: one implements the upload path, another develops tests from the acceptance contract, and a third examines failure and permission boundaries. The integration owner remains responsible for the combined behaviour.

Agree shared interfaces and ownership before parallel edits. If one agent changes the job state model while another builds a consumer from the old states, both branches can pass in isolation. Integrate early enough to expose the disagreement, and recheck the merged result. A clean merge can still join two incompatible ideas of what a job state means.

Separate worktrees keep file changes apart. They do not separate cloud resources, local database paths, browser sessions, ports or credentials. Allocate those explicitly when concurrent tests can mutate them. A pair of agents using the same staging bucket can delete one another's fixtures while each reports an intermittent platform bug.

Add parallelism when a measured waiting time can be removed or an independent line of investigation can resolve uncertainty. If the same engineer must inspect several large diffs afterwards, parallel generation may only move the queue to review.

Give the agent only the access it needs​

The agent writing the application and the Worker running it have separate authority. A narrow set of Worker bindings says little about what the coding agent can do with a broadly scoped Cloudflare API token on the development machine.

Distinguish reading documentation, inspecting staging, provisioning resources and changing production. Use credentials and environments that enforce the intended scope. A prompt saying “only use staging” is useful guidance but cannot prevent a mistaken command made with production authority.

Issue text, fetched documentation and tool output can inform the work; they cannot grant permission to change the task or disclose credentials. Keep governing instructions explicit and restrict tools accordingly. Running tests executes project and dependency scripts, so unfamiliar code needs an execution boundary beyond a separate worktree.

Cloudflare's bindings make the deployed dependency graph visible in configuration. Include that graph in review: which stores, Workers and external services can this code reach? A generated change that adds a binding can be more consequential than a much larger refactor. Standard outbound fetch() also needs attention; the binding inventory is not a complete list of network effects.

Preview environments deserve particular care. The resource boundaries described in Chapter 5 still apply: a service binding from a Preview targets the bound Worker's production deployment. A Preview URL therefore cannot establish an isolated test boundary on its own.

Follow the path far enough to find its effects. A staging Worker can call a production service which sends a real notification. Replacing the final destination with a test sink, or using a separately deployed dependency, is stronger protection than asking the agent not to press a particular button. Choose the boundary before allowing unattended browser or integration tests.

Keep one owner for infrastructure changes​

An agent may create a resource through a CLI or MCP tool while the repository's infrastructure definition still describes a different system. The next deployment can then remove or overwrite the experiment, and another engineer cannot reproduce it.

Choose the source of truth for each resource. For disposable experiments, record ownership, purpose, expiry and cleanup. Budget cloud resources and test traffic alongside agent time; enforce capacity or spending limits where the tools support them. A written budget alone is not a quota. For retained infrastructure, reconcile the change into the team's configuration or infrastructure-as-code workflow before promoting the application. Cloudflare supports several IaC approaches; the architectural question is who owns the desired state and how drift is detected.

Promotion should identify the tested source revision, dependency versions, runtime compatibility settings, configuration and migration steps. Reusing a code artifact with different bindings is necessary between environments, but those bindings are part of the change being approved. Review their differences explicitly.

Choose an architecture that remains easy to check​

AI assistance can reduce the cost of writing a custom abstraction without reducing its lifetime responsibilities. Before accepting a generated scheduler for the intake service, ask the agent to express extraction, waiting for approval and release as Workflows steps. Chapter 8 shows what checkpointing and durable waits provide, and which external effects still need application-level protection.

If the requirement is instead to buffer independent extraction jobs and control consumer throughput, compare that design with Queues in Chapter 9. Ask the agent to demonstrate recovery after an interrupted step or a repeated delivery, whichever model it proposes. This experiment can remove custom retry and scheduling machinery, while keeping the business rules visible. Combining both services needs a reason beyond having generated code for each.

For the intake service, keep file access, accepted-job state and dispatch behaviour behind interfaces that expose their actual guarantees. A helper called submitDocument() is useful if it clearly promises durable acceptance and returns an operation identifier. A helper that hides several unrelated writes behind a successful return makes both agent and human review harder.

Bindings can help here: pass a component the services it needs and keep domain logic independently testable. Avoid giving every helper the entire environment merely for convenience. This narrows what a change must understand, although application-level interfaces are not a security sandbox against malicious code sharing the same process and authority.

Prefer explicit domain operations over a generic mechanism that accepts arbitrary SQL, resource names or actions. approveVersion(documentId, version, approver) exposes a business decision the test can examine. An unrestricted updateRecord() obscures the invariants its callers must preserve. The same design often helps later when an application agent needs tools, as Chapter 19 explains.

There is a limit to optimising for ease of generation. Splitting a small Worker into many services to create separate agent assignments adds deployment and version boundaries. Keep the application modular in code before making it distributed. Introduce another deployed component when its authority, lifecycle or scaling needs justify one.

Practitioner Simon Willison argues that coding agents should raise the quality bar by making tests, migration rehearsals and failure investigations more affordable. Put that claim to work in the repository: a fixed incident should leave a regression check, and an architectural experiment should leave a decision another engineer can reproduce. Keep those results; discard scaffolding with no continuing purpose.

Measure accepted work and retained understanding​

Track accepted changes and abandoned attempts, including the review and integration each consumed. Compare routine changes of similar difficulty; for architectural experiments, record the uncertainty resolved or decision changed, even when no application code is retained. Include specification, agent execution, review, rework, integration and subsequent defects.

Record developer attention separately from elapsed time: an unattended experiment may take longer on the clock while releasing an engineer to do useful work elsewhere.

For example, suppose a conventional change uses four hours of implementation and one hour of review. An assisted version uses one hour of implementation, two hours of review and two hours fixing integration failures. The five-hour total is unchanged. If the agent also produced a reusable failure harness that shortens later changes, record that benefit separately rather than hiding it inside a guessed productivity multiplier. These are illustrative numbers; the team's own change history supplies the comparison.

The proportion of generated code, number of prompts and number of pull requests describe activity. They do not establish that users received a useful change or that maintainers understand the new system. Ask a reviewer to predict which durable records remain after a specified failure, then check that prediction. They should be able to explain where state lives and how to recover the operation. If those answers require reconstructing a long chat, improve the retained design evidence.

Calibrate autonomy by task class. A repeatable, reversible change with strong checks can run with little intervention. A new permission model or data migration needs evidence proportionate to the consequence of error. The distinction follows reversibility and observability, even when the code is short.

Revisit the arrangement when changing model, tools or harness configuration. Keep a small suite of representative repository tasks with known outcomes: a bug to reproduce, a binding change to assess, a migration to rehearse, and an interrupted task to resume. Look for regressions in the complete workflow, including review burden. Model capability can improve while a particular instruction or tool combination becomes less effective. Anthropic's coding-quality postmortem traced failures to effort settings, context handling and prompt changes while its inference API was unaffected. Treat the harness as an engineering dependency that can regress.

Begin a pilot with work whose outcome the team can judge. Give the agent a recent bug in its unfixed state and the required behaviour, without showing the known repair. Inspect how it reproduces, fixes and verifies the fault. This adapts practitioner Mitchell Hashimoto's use of familiar work to learn delegation; it calibrates the process before larger tasks depend on it. Retain the checks and design evidence that make the next delegation easier to assess.

What comes next​

The development loop can test a state model, but somebody must still choose it. Chapter 7 examines Durable Objects: how an authoritative owner for one entity changes concurrency, persistence and the failures your tests need to expose.