[{"data":1,"prerenderedAt":1041},["ShallowReactive",2],{"article_list_spec-driven development_":3},[4],{"_path":5,"_dir":6,"_draft":7,"_partial":7,"_locale":8,"title":9,"description":10,"publishDate":11,"image":12,"author":13,"tags":16,"excerpt":10,"body":23,"_type":1035,"_id":1036,"_source":1037,"_file":1038,"_stem":1039,"_extension":1040},"/ewahl/2026-08/ai-coding-agents-software-methodology","2026-08",false,"","AI Agents Didn't Invent a New Way to Build Software. They Funded the Old One.","The right way to build software has been known since 1970, proven on flagship programs, and priced out of reach of everyone else ever since. This is the story of that fifty-year gap — what the theory prescribed, what real-world budgets actually permitted, and how AI coding agents are closing it: one old practice at a time.","2026-08-25","/ewahl/2026-08/img/ai-coding-agents-software-methodology.jpg",{"name":14,"user":15},"Edward F. Wahl","ewahl",[17,18,19,20,21,22],"ai","agentic ai","ai coding agents","humans in the loop","spec-driven development","agile software development",{"type":24,"children":25,"toc":998},"root",[26,37,41,46,63,68,73,92,104,116,123,130,156,162,174,180,185,191,208,213,218,224,230,249,255,267,273,285,291,310,316,328,334,346,352,357,362,374,379,391,397,423,429,455,481,514,520,546,552,557,562,592,598,616,622,627,633,651,657,662,669,674,679,691,696,702,707,713,718,744,750,762,767,772,777,780,786,803,813,823,833,843,853,863,873,879],{"type":27,"tag":28,"props":29,"children":30},"element","p",{},[31],{"type":27,"tag":32,"props":33,"children":34},"em",{},[35],{"type":36,"value":10},"text",{"type":27,"tag":38,"props":39,"children":40},"hr",{},[],{"type":27,"tag":28,"props":42,"children":43},{},[44],{"type":36,"value":45},"I run engineering at Art+Logic, a custom software firm, which means I watch a lot of projects start. For twenty years they started more or less the same way — different clients, same first act. Not anymore. Lately, the first two weeks of a greenfield build go like this.",{"type":27,"tag":28,"props":47,"children":48},{},[49,51,56,58],{"type":36,"value":50},"The client talks. We ask questions — the annoying kind, about edge cases and exceptions and what happens when two rules collide. A few days later, not months, they click through a working prototype of the thing they described. It has their workflows in it, their vocabulary, their weird approval chain. We ask two questions. The first is ",{"type":27,"tag":32,"props":52,"children":53},{},[54],{"type":36,"value":55},"is this what you meant?",{"type":36,"value":57}," — the old requirements question, answerable now by pointing instead of by interpreting a document. The second matters more: ",{"type":27,"tag":32,"props":59,"children":60},{},[61],{"type":36,"value":62},"now that you can see what you asked for — does it hold up?",{"type":27,"tag":28,"props":64,"children":65},{},[66],{"type":36,"value":67},"Surprisingly often, it doesn't. The client looks at the screen and discovers that the rule they stated in the meeting, the one everyone nodded at, produces something nobody wants. Not a miscommunication — we built what they said. A discovery. Then the prototype gets thrown away. Its code was never going to ship; it existed to carry a conversation that documents alone couldn't carry. What accumulates instead, quietly, are the durable artifacts: the requirements the prototypes falsified their way toward, the standards, the scenarios that will hold the real system to account.",{"type":27,"tag":28,"props":69,"children":70},{},[71],{"type":36,"value":72},"From the outside this reads as the bleeding edge. There's an AI agent under every part of it. It is also, almost line for line, the process the field's founders started prescribing in the 1970s and '80s — and the part that should reframe your whole view of the current moment is this: that process was never hypothetical. It ran, for decades, at the top of the market. David Parnas ran his \"rational design process\" on the A-7E aircraft's operational flight software for the Naval Research Laboratory. And across the street from the Johnson Space Center, the on-board shuttle group (260 people maintaining the space shuttle's 420,000-line flight software, profiled by Charles Fishman in 1996 as \"They Write the Right Stuff\") ran the full program: roughly a third of the work done before anyone wrote a line of code, no line changed without an agreed change to the specification first. The results matched: the last three releases had one error each, against an estimated five thousand for commercial software of equivalent complexity. The price was part of the profile too: $35 million a year, for one program that flies one spaceship, in a shop where, as Fishman put it, \"money is not the critical constraint.\"",{"type":27,"tag":28,"props":74,"children":75},{},[76,78,83,85,90],{"type":36,"value":77},"One word in that paragraph is going to carry the whole essay, so let's define it now. A ",{"type":27,"tag":32,"props":79,"children":80},{},[81],{"type":36,"value":82},"specification",{"type":36,"value":84}," — the spec — is the complete statement of what a system is supposed to do: not \"handle invoices\" but every rule, every exception, every edge case, written down precisely enough that two people reading it would build the same behavior. The spec is not the code. The code is what a computer executes; the spec is what the code is ",{"type":27,"tag":32,"props":86,"children":87},{},[88],{"type":36,"value":89},"for",{"type":36,"value":91},".",{"type":27,"tag":28,"props":93,"children":94},{},[95,97,102],{"type":36,"value":96},"So: the right way to build software was known, published, and demonstrated within living memory of the moon landings. What it never was, for anyone below the flagship tier, was ",{"type":27,"tag":32,"props":98,"children":99},{},[100],{"type":36,"value":101},"affordable",{"type":36,"value":103},". Everything else in the industry's fifty-year methodology history — waterfall, agile, and the wars between them — is best understood as a series of compromises with that price tag, each one keeping the part of the program it could pay for and dropping the rest.",{"type":27,"tag":28,"props":105,"children":106},{},[107,109,114],{"type":36,"value":108},"AI coding agents have not invented a new way to build software. They are collapsing the price of the old way — practice by practice. So let's take the old program apart into its practices and run each one through the same three questions: what did the theory say, what could the real world actually afford, and what does an agentic build do instead? Then we'll put it back together and walk a project end to end. (And if the word \"specification\" is already making you think ",{"type":27,"tag":32,"props":110,"children":111},{},[112],{"type":36,"value":113},"big rigid plan, signed in blood, wrong by June",{"type":36,"value":115}," — good instinct. Hold it. Adapting to change is half of this story, and it survives.)",{"type":27,"tag":117,"props":118,"children":120},"h2",{"id":119},"the-throwaway-you-dont-know-what-to-build-until-youve-built-it",[121],{"type":36,"value":122},"The throwaway: you don't know what to build until you've built it",{"type":27,"tag":124,"props":125,"children":127},"h3",{"id":126},"royces-warning-and-the-diagram-everyone-misread",[128],{"type":36,"value":129},"Royce's warning and the diagram everyone misread",{"type":27,"tag":28,"props":131,"children":132},{},[133,135,140,142,147,149,154],{"type":36,"value":134},"Start with the most misread paper in the discipline. In 1970 an engineer named Winston Royce published \"Managing the Development of Large Software Systems,\" and it contains a diagram you have effectively seen even if you've never opened a software methodology book: a series of boxes cascading downhill — gather the requirements, then design the system, then write the code, then test it, then deliver — each stage finished and signed off before the next begins. The cascade look is where the industry's nickname for this process came from: ",{"type":27,"tag":32,"props":136,"children":137},{},[138],{"type":36,"value":139},"waterfall",{"type":36,"value":141},". Royce's paper is universally cited as waterfall's origin. The joke history played on him: the word \"waterfall\" appears nowhere in it, and the diagram appears as a ",{"type":27,"tag":32,"props":143,"children":144},{},[145],{"type":36,"value":146},"warning",{"type":36,"value":148},". Directly beneath it he wrote, \"I believe in this concept, but the implementation described above is risky and invites failure.\" His reasoning was simple and turned out to be prophetic: in a single downhill pass, the first honest feedback — real users touching real software — arrives at the very end, after all the money is spent, which is the most expensive possible moment to learn something. What Royce actually prescribed was the opposite of what got built in his name: ",{"type":27,"tag":32,"props":150,"children":151},{},[152],{"type":36,"value":153},"do it twice",{"type":36,"value":155},", with a preliminary pilot version built specifically to be discarded once it had taught you where the trouble lived, and involve the customer throughout, because the alternative is discovering at delivery what they meant at kickoff.",{"type":27,"tag":124,"props":157,"children":159},{"id":158},"brooks-plan-to-throw-one-away",[160],{"type":36,"value":161},"Brooks: plan to throw one away",{"type":27,"tag":28,"props":163,"children":164},{},[165,167,172],{"type":36,"value":166},"Fred Brooks (manager of IBM's gigantic System/360 software effort in the 1960s, and the closest thing the field has to a founding elder) made the same point unforgettable in his 1975 book ",{"type":27,"tag":32,"props":168,"children":169},{},[170],{"type":36,"value":171},"The Mythical Man-Month",{"type":36,"value":173},": \"plan to throw one away; you will, anyhow.\" His argument was an economist's, not an idealist's. The first system built is the one on which you learn what the system should have been. That's unavoidable. So the only real choice is whether to budget for the throwaway or to ship it to paying customers and call the learning process a maintenance contract. By 1986, in his essay \"No Silver Bullet,\" Brooks had upgraded the pilot from good practice to the builder's core duty: since clients cannot fully know what they want until they see something running, the single most important service a software builder performs is the iterative extraction and refinement of the requirements, through rapid prototyping. The prototype, in Brooks's telling, is not a demo. It is a requirements instrument.",{"type":27,"tag":124,"props":175,"children":177},{"id":176},"what-the-throwaway-actually-cost",[178],{"type":36,"value":179},"What the throwaway actually cost",{"type":27,"tag":28,"props":181,"children":182},{},[183],{"type":36,"value":184},"Now price it. Building a system twice, at human implementation rates, was a proposal that died in every budget meeting it ever entered — so the real world split into the two outcomes Brooks predicted. Most projects skipped the pilot entirely and ran the waterfall Royce warned about. If you have ever been the client on one of these, you know the shape from the inside: months of meetings distilled into a fat requirements document, a signature page, a long quiet stretch of invoices, and then a delivery date on which you finally touch the system and begin composing the sentence \"this isn't quite what we meant,\" with the budget already spent and every correction now billed as a change order. The requirements were frozen at the moment of least knowledge, which is the beginning, and nothing learned along the way had anywhere to go. The rest of the projects built the pilot and then committed the sin Brooks named specifically: they shipped it. Of course they did. Once a demo exists, building the real system \"again\" is an expense nobody will authorize, so demo code became production code, shortcuts and all. Even Brooks retreated. Twenty years later he called his own maxim \"wrong, not because it is too radical, but because it is too simplistic,\" his stated reason being that a full up-front throwaway assumed waterfall economics. The father of the idea didn't back away because it was wrong. He backed away because of what it cost.",{"type":27,"tag":124,"props":186,"children":188},{"id":187},"the-throwaway-on-an-agentic-build",[189],{"type":36,"value":190},"The throwaway on an agentic build",{"type":27,"tag":28,"props":192,"children":193},{},[194,196,201,203],{"type":36,"value":195},"On an agentic build, the cost is gone, and the practice comes back exactly as specified — so let me make it concrete with an example I'll keep returning to for the rest of this essay. It's hypothetical, but every business has a rule shaped like it. Suppose that in a requirements meeting, the client's CFO states a policy: ",{"type":27,"tag":32,"props":197,"children":198},{},[199],{"type":36,"value":200},"any invoice over $10,000 requires a second approval from a manager.",{"type":36,"value":202}," Everyone nods. It goes in the document. In the old world, that sentence gets implemented in month four and questioned by nobody. In ours, it's running in a prototype within days, and the client clicks through it — at which point someone from billing points at the screen and says: wait. Our biggest customer pays a $14,000 retainer on the first of every month. As written, this rule routes our single most routine transaction through a manager's inbox, twelve times a year, forever. ",{"type":27,"tag":32,"props":204,"children":205},{},[206],{"type":36,"value":207},"That's not what we meant. We meant unusual invoices.",{"type":27,"tag":28,"props":209,"children":210},{},[211],{"type":36,"value":212},"Notice what just happened. Our first question — is this what you meant? — would catch a misunderstanding, but there was no misunderstanding; we built the stated rule faithfully. It was the second question — now that you can see it, does it hold up? — that caught the real problem: the stated rule itself was wrong, and it took a running system to make that visible. That defect just got fixed for the cost of a conversation. In a waterfall project it would have been discovered in production, by an annoyed manager, and fixed as a change order.",{"type":27,"tag":28,"props":214,"children":215},{},[216],{"type":36,"value":217},"And when the conversation moves on, the prototype gets deleted — which is the discipline that was economically impossible for fifty years. When a rebuild costs less than the meeting where you'd argue about keeping it, the prototype can finally be what Royce and Brooks wanted it to be: an instrument, used and retired. None of its code ships. What ships forward is what it taught the spec. Brooks's dictum that clients can't know what they want until they see something running was always true. What was missing was the price.",{"type":27,"tag":117,"props":219,"children":221},{"id":220},"the-living-document-writing-it-down-and-keeping-it-from-becoming-a-lie",[222],{"type":36,"value":223},"The living document: writing it down, and keeping it from becoming a lie",{"type":27,"tag":124,"props":225,"children":227},{"id":226},"essential-and-accidental-complexity",[228],{"type":36,"value":229},"Essential and accidental complexity",{"type":27,"tag":28,"props":231,"children":232},{},[233,235,240,242,247],{"type":36,"value":234},"The theory of why you write the specification down — and why writing it down was never the hard part — has two layers, and the deeper one is the less famous. The famous layer is Brooks again. In \"No Silver Bullet\" he split software's difficulties into two kinds. The ",{"type":27,"tag":32,"props":236,"children":237},{},[238],{"type":36,"value":239},"accidental",{"type":36,"value":241}," difficulties are the friction of expression — the labor of translating a decision into working code, wrestling with languages and tools and plumbing. The ",{"type":27,"tag":32,"props":243,"children":244},{},[245],{"type":36,"value":246},"essential",{"type":36,"value":248}," difficulty is the conceptual construct itself: deciding what the rules actually are, how they interact, what happens in every edge case. In our running example, the accidental part is making the approval screen exist. The essential part is knowing that the threshold should exempt routine retainers — and every decision like it, across the whole system, held consistent. Brooks's verdict has been reprinted for four decades: \"The hardest single part of building a software system is deciding precisely what to build.\" Nothing cripples a system like getting that wrong, and nothing is costlier to fix late.",{"type":27,"tag":124,"props":250,"children":252},{"id":251},"parnas-how-and-why-to-fake-it",[253],{"type":36,"value":254},"Parnas: how and why to fake it",{"type":27,"tag":28,"props":256,"children":257},{},[258,260,265],{"type":36,"value":259},"The less famous layer arrived the same year, in the most honest paper in software engineering. The honesty is in the title: \"A Rational Design Process: How and Why to Fake It,\" by David Parnas and Paul Clements. The ideal process, they conceded, is one no real project follows: every design decision derived from a precise, documented statement of what the system must do, every choice traceable to a reason. Humans don't work that way and schedules don't permit it. Their fallback: follow the ideal as closely as you can, and ",{"type":27,"tag":32,"props":261,"children":262},{},[263],{"type":36,"value":264},"produce the documentation the ideal process would have produced",{"type":36,"value":266}," — even though you didn't live it — because those documents are what make a system reviewable, maintainable, and comprehensible to the next person who has to touch it. The field's considered response to \"the right way costs more than we have\" was to institutionalize forging the paper trail.",{"type":27,"tag":124,"props":268,"children":270},{"id":269},"the-confident-lie",[271],{"type":36,"value":272},"The confident lie",{"type":27,"tag":28,"props":274,"children":275},{},[276,278,283],{"type":36,"value":277},"What the real world could afford was less than that, and the reason is a property of documentation that makes half-funding it worse than not funding it at all. Software changes constantly. Documents don't change themselves. And a document that has drifted from the code is not neutral — it is a confident lie. (Every company has one: the wiki page, last edited three years ago, that states the approval threshold is $5,000. It hasn't been $5,000 since the reorg. New employees read it anyway.) Keeping documentation ",{"type":27,"tag":32,"props":279,"children":280},{},[281],{"type":36,"value":282},"true",{"type":36,"value":284}," — re-verifying its correspondence with a moving system, change after change — was labor nobody could bill for. So the industry retreated twice. The first retreat was Parnas's own advice: fake it. The second was more radical, more successful, and deserves a proper introduction, because it reshaped the entire industry.",{"type":27,"tag":124,"props":286,"children":288},{"id":287},"agiles-bargain-the-software-became-the-document",[289],{"type":36,"value":290},"Agile's bargain: the software became the document",{"type":27,"tag":28,"props":292,"children":293},{},[294,296,301,303,308],{"type":36,"value":295},"In 2001, seventeen veteran practitioners met at a ski lodge in Utah and published a short statement of values called the Manifesto for Agile Software Development. \"Agile\" is the name that stuck for the whole family of methods it blessed, and if you've worked anywhere near a modern software team you've lived inside its vocabulary: instead of one big plan executed over a year, work proceeds in short cycles — ",{"type":27,"tag":32,"props":297,"children":298},{},[299],{"type":36,"value":300},"sprints",{"type":36,"value":302},", typically two weeks — each of which delivers a small, working, usable slice of the system; the client sees running software every cycle; priorities get re-decided between cycles based on what everyone just learned. It was, in large part, the industry's immune response to waterfall — a way to stop freezing requirements at the moment of least knowledge. And among its declared values was this one: \"working software over comprehensive documentation,\" with the explicit caveat that the items on the right had value too, just less. Read that as an economic judgment and it was exactly right for its decade. If you cannot afford to keep the documents true, then a team forced to choose should trust the artifact that provably runs. So agile demoted the artifact that couldn't be kept in sync and promoted the one that could be executed. ",{"type":27,"tag":32,"props":304,"children":305},{},[306],{"type":36,"value":307},"The software became the document.",{"type":36,"value":309}," Want to know what the approval rule really is today? Don't read the wiki — it lies. Read the code, or find the person who last changed it and ask.",{"type":27,"tag":124,"props":311,"children":313},{"id":312},"naur-the-program-lives-in-peoples-heads",[314],{"type":36,"value":315},"Naur: the program lives in people's heads",{"type":27,"tag":28,"props":317,"children":318},{},[319,321,326],{"type":36,"value":320},"Which brings us to the endpoint of that retreat, described with eerie precision back in 1985 by a Danish computer scientist named Peter Naur, in an essay called \"Programming as Theory Building.\" A program, Naur argued, is not really its source text. It is a ",{"type":27,"tag":32,"props":322,"children":323},{},[324],{"type":36,"value":325},"theory",{"type":36,"value":327}," — a live understanding, held by people, of how the business's problems and the code's structure correspond — and the essential part of that understanding \"could not conceivably be expressed, but is inextricably bound to human beings.\" You have met Naur's theory even if you've never read him. It's why the answer to \"why does the discount field cap at 15 percent?\" is not in any document but in the memory of whoever met that one customer in 2019. It's why, when the engineer who built the billing system leaves, the billing system doesn't stop running but does, in a real sense, die. It can still be executed. It can no longer be confidently changed. For a custom software firm, this was always the uncomfortable truth of the balance sheet: the asset was never the code. It was the theory in the senior engineers' heads, and it walked out the door every evening.",{"type":27,"tag":124,"props":329,"children":331},{"id":330},"the-first-reliable-reader",[332],{"type":36,"value":333},"The first reliable reader",{"type":27,"tag":28,"props":335,"children":336},{},[337,339,344],{"type":36,"value":338},"An agentic build reverses the retreat, for a bluntly mechanical reason: an AI agent cannot read anyone's head. It consumes artifacts — requirements, standards, written scenarios — and it produces exactly what the artifacts say, which means everything that stays tacit is now, operationally, not part of the project. That sounds like a burden, and then you notice what it buys. The agent is the first teammate that consumes documentation as an operational input rather than honoring it as a virtuous intention — the first reader the documents have ever reliably had. Humans could always substitute a hallway conversation for the wiki, which is why the wiki died; an agent has no hallway. Parnas told the field to fake the rational process because living it was unaffordable. That excuse just expired. So on our projects the specification becomes a first-class artifact and stays one: requirements distilled from the actual mass of meetings, emails, and threads — an agent can hold all of it at once and surface where March's decision quietly contradicts May's. Add the technical standards, stated up front by a senior engineer, defining the right way to build on the chosen foundations. All of it versioned like code, enforced like code, and edited for the life of the system. What does ",{"type":27,"tag":32,"props":340,"children":341},{},[342],{"type":36,"value":343},"not",{"type":36,"value":345}," reprice is the essential difficulty itself. Deciding precisely what to build — that retainers are routine, that refunds are different, that Canada is coming — costs what it cost in 1986, because it was never a typing problem. And on a project where the typing is nearly free, it is most of the remaining bill.",{"type":27,"tag":117,"props":347,"children":349},{"id":348},"the-reconciliation-every-change-checked-against-the-whole",[350],{"type":36,"value":351},"The reconciliation: every change checked against the whole",{"type":27,"tag":28,"props":353,"children":354},{},[355],{"type":36,"value":356},"The purest statement of this practice is the shuttle group's standing rule: no coder changes a line without an agreed change to the specification first, and no spec change lands without both sides understanding everything it touches. Every change, reconciled against everything the system is supposed to be, before it exists. That is what \"a third of the process happens before code\" actually means in practice. And the theory of why ordinary organizations can't do this was written by Brooks too — it is the argument his book's title is famous for. The \"mythical man-month\" is the fallacy that people and time are interchangeable, that a late project can be rescued by adding staff; Brooks's Law says the opposite happens, and the arithmetic behind it is brutal in an ordinary way. Three people on a project have three lines of communication to maintain. Ten people have forty-five. Every person you add to buy more review capacity adds coordination overhead faster than they add review — which is why \"check every change against everything\" is not a thing you can hire your way into.",{"type":27,"tag":28,"props":358,"children":359},{},[360],{"type":36,"value":361},"What that scarcity cost in practice is the untold half of the agile story. Agile's bet — the code is the document, each sprint makes a change, you learn by shipping — is powerful. It also carries a cost curve that every team that has ridden it for a few years knows in their bones, even if nobody put it on a slide. As the system grows, the fraction of the big picture any one human can hold in their head while reviewing a change shrinks.",{"type":27,"tag":28,"props":363,"children":364},{},[365,367,372],{"type":36,"value":366},"Return to our invoice rule for the concrete version. Two years into the system's life, a developer adds a \"rush order\" feature: a streamlined screen that creates invoices in three clicks for time-sensitive deals. The feature is well built. Its tests pass. Sales loves it. And nobody in the review noticed that invoices created through the new screen never pass through the approval check, because the approval logic was wired to the old screen, and the person who built that wiring left last spring. The change was ",{"type":27,"tag":32,"props":368,"children":369},{},[370],{"type":36,"value":371},"locally correct and globally wrong",{"type":36,"value":373}," — perfect in its own corner, a policy violation in the whole — and no individual did anything unreasonable; the knowledge needed to catch it was simply spread across more heads and more code than any reviewer could hold.",{"type":27,"tag":28,"props":375,"children":376},{},[377],{"type":36,"value":378},"That is the quiet tax of agile at scale. Regressions get subtler. Estimates grow multipliers. \"Small change\" stops being a coherent category. Not because agile is wrong, but because its source of truth is the code — an artifact optimized for a computer to execute, not for anyone to check a change against everything the system is supposed to be. The pragmatic ceiling on agile efficiency was always the cost of that check: the same scarce input, review, that made living documentation unaffordable in the first place.",{"type":27,"tag":28,"props":380,"children":381},{},[382,384,389],{"type":36,"value":383},"This is where agents change something structural, because the capability that matters most is not writing code. It is reading everything, every time. An agent can hold the requirements, the standards, and the implementation in view at once and evaluate a proposed change against all three — every change, at a unit cost that makes whole-system review economically routine. The shuttle group's rule, as a default setting. Run the rush-order story again under that regime: the change arrives, and the reconciliation flags it — ",{"type":27,"tag":32,"props":385,"children":386},{},[387],{"type":36,"value":388},"this creates invoices without invoking the approval requirement stated in the spec",{"type":36,"value":390}," — before it lands, because something actually compared the change against the whole statement of what the system is supposed to be, which is exactly the review no human team could afford to perform on every change. So the agile loop keeps running — sprints, changes, learning by shipping — but each change is now reconciled against both the specification and the codebase as it lands. When implementation teaches that a requirement was wrong (the retainer discovery, at any point in the project's life), the requirement gets updated, not silently overruled by the code. When a change is locally green and globally wrong, the breadth of review that used to be unaffordable is what stands positioned to catch it. Specification and system stay in sync, and the sync is re-earned at every change. Which is why this is not waterfall wearing new clothes: the spec is not frozen at kickoff — on our projects it is the most-edited artifact in the repository, and the arrow between spec and code runs in both directions for the life of the system. Iteration survived. What died was the forced choice between iterating and staying specified.",{"type":27,"tag":117,"props":392,"children":394},{"id":393},"the-executable-fraction-what-a-test-suite-actually-is",[395],{"type":36,"value":396},"The executable fraction: what a test suite actually is",{"type":27,"tag":28,"props":398,"children":399},{},[400,402,407,409,414,416,421],{"type":36,"value":401},"One more piece of vocabulary, because the last practice lives inside it. A ",{"type":27,"tag":32,"props":403,"children":404},{},[405],{"type":36,"value":406},"test",{"type":36,"value":408},", in software, is a small program whose job is to use your system and check the answer. For the invoice rule, one test might create a $12,000 invoice and verify the system demands a second approval. Another creates a $9,000 invoice and verifies it doesn't. A real project accumulates hundreds or thousands of these — the ",{"type":27,"tag":32,"props":410,"children":411},{},[412],{"type":36,"value":413},"test suite",{"type":36,"value":415}," — and runs the entire suite automatically every time anything changes. When every test passes, the dashboard glows green; \"the build is green\" is industry shorthand for ",{"type":27,"tag":32,"props":417,"children":418},{},[419],{"type":36,"value":420},"all checks pass",{"type":36,"value":422},". Tests are the software profession's answer to \"how do you change one thing without breaking everything else.\" They are as close to a superpower as the field has.",{"type":27,"tag":124,"props":424,"children":426},{"id":425},"what-bdd-and-specification-by-example-got-right",[427],{"type":36,"value":428},"What BDD and specification by example got right",{"type":27,"tag":28,"props":430,"children":431},{},[432,434,439,441,446,448,453],{"type":36,"value":433},"The theory arrived two decades ago and was undersold by its own modest framing. In 2006 a consultant named Dan North published \"Introducing BDD\" (behavior-driven development), arguing that tests should be written as descriptions of behavior that a businessperson could read. The style that grew out of it expresses each scenario in three moves — ",{"type":27,"tag":32,"props":435,"children":436},{},[437],{"type":36,"value":438},"Given",{"type":36,"value":440}," a situation, ",{"type":27,"tag":32,"props":442,"children":443},{},[444],{"type":36,"value":445},"When",{"type":36,"value":447}," something happens, ",{"type":27,"tag":32,"props":449,"children":450},{},[451],{"type":36,"value":452},"Then",{"type":36,"value":454}," here's what must be true — and it reads like this:",{"type":27,"tag":456,"props":457,"children":458},"blockquote",{},[459],{"type":27,"tag":28,"props":460,"children":461},{},[462,467,469,473,475,479],{"type":27,"tag":463,"props":464,"children":465},"strong",{},[466],{"type":36,"value":438},{"type":36,"value":468}," a customer on the preferred list with a monthly retainer\n",{"type":27,"tag":463,"props":470,"children":471},{},[472],{"type":36,"value":445},{"type":36,"value":474}," an invoice over $10,000 is created for that retainer\n",{"type":27,"tag":463,"props":476,"children":477},{},[478],{"type":36,"value":452},{"type":36,"value":480}," no second approval is required",{"type":27,"tag":28,"props":482,"children":483},{},[484,486,491,493,498,500,505,507,512],{"type":36,"value":485},"Gojko Adzic's book named the whole practice ",{"type":27,"tag":32,"props":487,"children":488},{},[489],{"type":36,"value":490},"Specification by Example",{"type":36,"value":492},", and in that community a test suite is called an ",{"type":27,"tag":32,"props":494,"children":495},{},[496],{"type":36,"value":497},"executable specification",{"type":36,"value":499},". The name is exactly right, and it is the bridge this essay has been building toward: a test ",{"type":27,"tag":463,"props":501,"children":502},{},[503],{"type":36,"value":504},"is",{"type":36,"value":506}," a specification — a piece of the spec, written so precisely that a computer can check it. It's just a ",{"type":27,"tag":32,"props":508,"children":509},{},[510],{"type":36,"value":511},"partial",{"type":36,"value":513}," one: the fraction of your intent that somebody managed to make executable. Call the fraction 40 percent. The number is invented; the existence of the remainder is not. Every suite asserts some things and is silent about the rest, and the rest is not small — the intent behind the assertions, the meaning of the domain's words, the edge behavior nobody wrote down because to a human it was obvious. Nobody wrote the test that says a credit memo — a refund — doesn't count toward the $10,000 threshold, because nobody in billing would ever dream of counting one.",{"type":27,"tag":124,"props":515,"children":517},{"id":516},"why-tdd-worked-the-developer-supplied-the-other-60-percent",[518],{"type":36,"value":519},"Why TDD worked: the developer supplied the other 60 percent",{"type":27,"tag":28,"props":521,"children":522},{},[523,525,530,532,537,539,544],{"type":36,"value":524},"Two things kept that gap from mattering, historically, and both are worth seeing clearly. The first is that the fraction was thinner than it had to be — comprehensive suites were among the first things budgets cut. The second is an accounting trick nobody noticed they were running, inside a practice called test-driven development, or TDD: the discipline, dominant among serious teams for two decades, of writing the test ",{"type":27,"tag":32,"props":526,"children":527},{},[528],{"type":36,"value":529},"before",{"type":36,"value":531}," the code — first the check, then the thing being checked, in tight little loops. TDD worked because the same head wrote the test and the code. The developer who wrote the assertions carried the other 60 percent across personally — of course refunds don't count; obviously voided invoices are excluded — honoring intent no test checked, because it was ",{"type":27,"tag":32,"props":533,"children":534},{},[535],{"type":36,"value":536},"their own understanding",{"type":36,"value":538},", and violating it would have felt like a bug even when nothing fired. The missing majority of the spec was supplied silently, from memory, on every loop. It never appeared in any artifact, so it never appeared in any budget. It looked free. And the size of what it was covering for turns out to be measurable. HumanEval (a benchmark, meaning a standardized exam for AI coding models: 164 small programming problems, each graded by fewer than ten tests) was for years the field's report card. In 2023 a project called EvalPlus (Liu et al., NeurIPS 2023) asked the obvious question: what happens if you grade the same answers against ",{"type":27,"tag":32,"props":540,"children":541},{},[542],{"type":36,"value":543},"more of the specification",{"type":36,"value":545},"? They generated eighty times as many tests per problem. Measured pass rates across twenty-six models fell by as much as 19 to 29 percent, model rankings flipped outright, and errors surfaced in some of HumanEval's own official solutions. Same models. Same answers. The only thing that changed was how much of the spec was written down.",{"type":27,"tag":124,"props":547,"children":549},{"id":548},"auditing-the-suite-held-out-scenarios-and-a-human-with-a-mouse",[550],{"type":36,"value":551},"Auditing the suite: held-out scenarios and a human with a mouse",{"type":27,"tag":28,"props":553,"children":554},{},[555],{"type":36,"value":556},"Agents end the old arrangement from both directions. An optimizer satisfies what you wrote, not what you meant — it will pass your 40 percent to the letter and route straight through the silence, not out of malice but because the silence is, to it, genuinely silent. Ask an agent to implement the approval rule as written and it will happily count that credit memo, because the artifact it read never said otherwise. TDD with an agent is not TDD, faster; it is your specification with its tacit half deleted and its explicit half satisfied exactly.",{"type":27,"tag":28,"props":558,"children":559},{},[560],{"type":36,"value":561},"So on an agentic build the suite gets the double treatment. It gets bigger, because the comprehensive suite, the one that exercises the whole application the way a real user would, is finally affordable. And it gets audited, because a richer visible suite is also a richer surface for an optimizer to satisfy in letter and defeat in spirit. The audit machinery has two parts. First, held-out scenarios: a second, private exam the agent never sees, run separately. If the visible suite is green and the hidden one is red, that divergence is an alarm. Second, a scheduled human exploratory pass: a person with a mouse and no stake in the suite's opinion, clicking around the way users actually will.",{"type":27,"tag":28,"props":563,"children":564},{},[565,567,572,574,579,581,590],{"type":36,"value":566},"And when something fails, the triage has exactly two honest exits. Either the requirement or the standard was wrong — in which case the ",{"type":27,"tag":32,"props":568,"children":569},{},[570],{"type":36,"value":571},"spec",{"type":36,"value":573}," improves and every future piece of work inherits the fix. Or the agent didn't follow a standard that was right — in which case the fix is to harden the ",{"type":27,"tag":32,"props":575,"children":576},{},[577],{"type":36,"value":578},"harness",{"type":36,"value":580},", the project's enforcement layer: automated checks that run on every single change and physically block the ones that violate a standard. The difference matters more than it sounds: a rule in a memo is a request; a rule in the harness is a locked door. What the fix is never allowed to be is a stern conversation with the agent about doing better next time — an agent's account of its own behavior is part of the output being optimized, and a safety mechanism the optimizer can produce is a safety mechanism made of output. House rule: the standard lives in the harness or it doesn't exist. (There is an adversarial version of this whole topic — what a strong optimizer does to a green suite when passing is the goal — and I've written about it separately: ",{"type":27,"tag":582,"props":583,"children":587},"a",{"href":584,"rel":585},"https://artandlogic.com/blog/ewahl/2026-07/goodharts-law-ai-coding-agents",[586],"nofollow",[588],{"type":36,"value":589},"goodharts law and ai coding agents",{"type":36,"value":591},".)",{"type":27,"tag":117,"props":593,"children":595},{"id":594},"the-build-end-to-end-five-phases-of-an-agentic-greenfield-project",[596],{"type":36,"value":597},"The build, end to end: five phases of an agentic greenfield project",{"type":27,"tag":28,"props":599,"children":600},{},[601,603,608,610,614],{"type":36,"value":602},"Step back and look at what those four practices — the throwaway prototype, the living document, whole-system reconciliation, and the executable specification — were before the agents. Waterfall kept the specification and lost the learning. Agile kept the learning and lost the specification. Each was the correct engineering economics for what review and rebuilding cost at the time — and the full program, living documents ",{"type":27,"tag":32,"props":604,"children":605},{},[606],{"type":36,"value":607},"and",{"type":36,"value":609}," continuous iteration ",{"type":27,"tag":32,"props":611,"children":612},{},[613],{"type":36,"value":607},{"type":36,"value":615}," whole-system verification, stayed where it had always been, on projects that could spend $35 million a year on 420,000 lines. Here is what it looks like reassembled, at normal-budget scale, as one project — five phases, including the one every methodology conversation forgets.",{"type":27,"tag":124,"props":617,"children":619},{"id":618},"discovery-prototypes-as-requirements-instruments",[620],{"type":36,"value":621},"Discovery: prototypes as requirements instruments",{"type":27,"tag":28,"props":623,"children":624},{},[625],{"type":36,"value":626},"The first weeks are the scene this essay opened with: conversations turned into disposable prototypes while they're still warm, the two questions asked and re-asked, requirements like the invoice rule falsified early and cheaply, the durable spec accumulating in the background. More running software than a traditional project would show in two months — and none of it the product. Each prototype is a question. The old week-two demo was a down payment: the code being applauded was the code the client would eventually own, with every shortcut taken to reach the applause compounding underneath it like debt.",{"type":27,"tag":124,"props":628,"children":630},{"id":629},"ramp-up-where-visible-progress-dips",[631],{"type":36,"value":632},"Ramp-up: where visible progress dips",{"type":27,"tag":28,"props":634,"children":635},{},[636,638,643,645,649],{"type":36,"value":637},"Then comes a compressed stretch, wedged between discovery's demos and implementation's releases, that I want to prepare you for honestly, because its visible progress dips exactly as its real progress accelerates. This is where the foundation goes in: the requirements formalized into scenario suites, the technical standards written down, the sample data, the harness — collectively, the documentation and validation guardrails that let an agent measure success while it builds. Nothing in this phase demos well, and somewhere in it every client asks some version of ",{"type":27,"tag":32,"props":639,"children":640},{},[641],{"type":36,"value":642},"when does the real application start?",{"type":36,"value":644}," The honest answer is that this ",{"type":27,"tag":32,"props":646,"children":647},{},[648],{"type":36,"value":504},{"type":36,"value":650}," the application — the part of it that determines whether everything after can be trusted. The clickable part earlier was the requirements conversation. This part is the definition of \"done\" being made checkable. It is also where the project's risk gets spent deliberately, at its cheapest. A requirement corrected here is an edit to a document; the same correction after launch is surgery on a live system with a year of data in it. That's not a research finding, it's arithmetic — but it runs against instinct, because the traditional arc trained everyone to read visible progress as real progress, and this is the one stretch where the two visibly part ways.",{"type":27,"tag":124,"props":652,"children":654},{"id":653},"implementation-agents-build-humans-hold-the-verdict",[655],{"type":36,"value":656},"Implementation: agents build, humans hold the verdict",{"type":27,"tag":28,"props":658,"children":659},{},[660],{"type":36,"value":661},"The bulk of the schedule, and the part everyone pictures when they picture AI coding — and it is genuinely great. The rhythm is recognizably agile: functionality specified against the standards in sprint-sized pieces, agents implementing, working software going out for client review on a frequent cadence. The demos are back, and this time they're the product. What's different is underneath. Every change is reconciled against the requirements, the standards, and the rest of the codebase before it lands — the rush-order check, running on everything — and whatever the build teaches flows back into the spec, so documentation, tests, and implementation stay in sync with each thing learned. Humans review the judgment calls the reconciliation surfaces rather than trying to re-derive that breadth themselves. The senior engineer's job here is not reading every line. It is recognizing mistakes while they are small — treating a suspicious success with the same attention as a failure, noticing that a hard piece went green a little too easily — and jumping on the discrepancy hard, because misunderstandings now compound at the speed of an agent, not the speed of a typist.",{"type":27,"tag":663,"props":664,"children":666},"h4",{"id":665},"the-requirement-nobody-wrote-down-a-validator-that-accepted-minus-five-hours",[667],{"type":36,"value":668},"The requirement nobody wrote down: a validator that accepted minus five hours",{"type":27,"tag":28,"props":670,"children":671},{},[672],{"type":36,"value":673},"I want to pause the walkthrough here for a story. Strictly an aside — the phases resume at the release — but it's the clearest picture I can give you of what that reviewer's job feels like from the inside.",{"type":27,"tag":28,"props":675,"children":676},{},[677],{"type":36,"value":678},"The rule the agent broke, on a recent build of ours, was that you cannot work negative hours. I want to be precise about how boring that is, because the boringness is the whole point. The ticket asked for a time entry — a worker records how long they spent on a task — and it specified, in writing, the one peculiar thing about how we count time: hours land on a 0.05-hour grid, three-minute increments, no odd fractions. The agent implemented that flawlessly. The validator it wrote checked divisibility by 0.05 and rejected anything off the grid — clean, correct, enforcing exactly the rule somebody had bothered to write down. It also cheerfully accepted -5.00, because minus five is a perfectly tidy multiple of 0.05; the tidiness check certified negative five hours as well-rounded and waved it through. Every test passed. Better than passed — the harness wasn't merely silent on the rule; it had written down the opposite and called it green. The test suite for that validator explicitly listed 0.00 as a valid value to accept.",{"type":27,"tag":28,"props":680,"children":681},{},[682,684,689],{"type":36,"value":683},"A human caught it, not a test, and not at first for the reason that mattered. The reviewer's note started small — a negative entry could corrupt the hours we mirror out to GitLab — and only widened, under a second look, into the actual hole: a negative row subtracts from the billable dollars in its rate group, drops silently out of the invoice preview so the work is simply never billed, and skews every report that sums hours. The post-mortem is where it got humbling. I went looking for the line the agent had violated, so I could say ",{"type":27,"tag":32,"props":685,"children":686},{},[687],{"type":36,"value":688},"there, that's the spec it ignored",{"type":36,"value":690}," — and there was no line. The requirement had never existed in writing. The closest artifact was a positivity constraint on a different table, the invoice line items, which proved only that we knew the rule well enough to enforce it wherever we'd happened to think of it, and never thought to write it on the table that actually stores the hours people work. You cannot blame the agent. It was faithful and it was competent; it did exactly, and only, what was written, and what was written was thorough about the exotic and mute about the obvious. It's the credit memo again: nobody in this shop would ever dream of logging minus five hours, and that is exactly why nobody ever wrote it down.",{"type":27,"tag":28,"props":692,"children":693},{},[694],{"type":36,"value":695},"The fix was to finally say the thing out loud, in the three places a machine can hear it: a check at the form boundary; a CHECK (time > 0) constraint in the database, so no code path can ever persist a negative hour again; and a property test in the invariant suite asserting that no zero or negative value is ever accepted — the same file where we had once parametrized 0.00 as fine, now inverted. That last diff is the scar I point to: a rule that lived for years below the waterline of writing, because it was too obvious to say, now pinned in a test that exists solely because a competent agent, reading a thorough spec, had no reason not to accept minus five hours. And this is the process working, not failing — the reviewer caught it early, by design. That is what the reviewer is for. But it is still humbling, because the miss was never in the 40 percent we argued over in the ticket. It was in the 60 percent nobody wrote down. And the 60 percent is made of things exactly this dumb.",{"type":27,"tag":124,"props":697,"children":699},{"id":698},"production-release-converging-on-the-specification",[700],{"type":36,"value":701},"Production release: converging on the specification",{"type":27,"tag":28,"props":703,"children":704},{},[705],{"type":36,"value":706},"Back to the walkthrough. By late implementation, something happens that took me a while to trust: the project converges on its specification. The requirements, the standards, and the scenario suites now describe the system that exists, and the harness enforces the correspondence — which changes what \"wrapping up\" means. In a traditional project this is where the deferred risk comes due: the integration crunch, the stabilization sprints, the feature freeze, the punch list that grows faster than it shrinks. Here, integration surprises have been getting caught continuously for months, so the endgame is short and quiet: hardening, cutover, and a final round of client adjustments — which are efficient for the same reason everything else has been, because a late change is an edit to the spec run through the harness, not an argument with a fragile codebase. Raise the approval threshold to $25,000 and exempt government customers? An edit, verified against everything, live before launch.",{"type":27,"tag":124,"props":708,"children":710},{"id":709},"maintenance-what-an-agent-built-system-costs-a-year-later",[711],{"type":36,"value":712},"Maintenance: what an agent-built system costs a year later",{"type":27,"tag":28,"props":714,"children":715},{},[716],{"type":36,"value":717},"And then the phase that every methodology debate skips, where a real and mostly invisible piece of the value sits. Consider what happens to a traditional project when active development ends. The team disperses to other work, and the unwritten 60 percent — the intent wedged between the tests and the code, the reasons living in people's heads — starts fading that same week. Come back in six or twelve months with a change request and the work is slower and riskier even if you get the same developers, because they must first re-learn their own system before they can safely change it. Usually you don't get the same developers, and then it's Naur's program death in slow motion: a system that still runs and can no longer be confidently changed. Now run the same clock against an agentic project. The knowledge that used to live in heads was encoded as the work happened — requirements, standards, scenario suites, and code, still in agreement because agreement was re-earned at every change. An agent picking the system up a year later reads the same artifacts and is exactly as fresh as the day active development paused. The project does not atrophy on the shelf. A routine change request in month fourteen costs what it cost in month four.",{"type":27,"tag":28,"props":719,"children":720},{},[721,723,728,730,735,737,742],{"type":36,"value":722},"Most of the time. Here is the caveat that keeps the maintenance claim honest. \"Just ask\" is true ",{"type":27,"tag":32,"props":724,"children":725},{},[726],{"type":36,"value":727},"inside the change-space the architecture anticipated",{"type":36,"value":729},". A change that violates a foundational assumption is a different animal: if the system was built assuming all money is in one currency and the client announces they're opening in Canada, that is a ",{"type":27,"tag":32,"props":731,"children":732},{},[733],{"type":36,"value":734},"refactor",{"type":36,"value":736}," (a rebuilding of internal structure, replumbing the house rather than repainting a room), and no harness converts it into a small ask, agents or no agents. What the approach buys is a shift in the cost ",{"type":27,"tag":32,"props":738,"children":739},{},[740],{"type":36,"value":741},"distribution",{"type":36,"value":743}," of change: the great mass of requests collapses toward trivial, and a fat tail of hard ones remains. Telling, quickly and correctly, which kind you're looking at is senior judgment, and it can't be delegated to the thing whose incentive is to say \"easy\" — which is why the senior engineer stands at the end of the project as well as the beginning, and \"the end,\" for a system that earns its keep, means years. And the deliverable underneath it all is the thing Naur said could never be handed off, or as close as I know how to get: not just a running system but its accumulated specification — explicit, versioned, enforced, still true. Naur is still right that some remainder lives only in heads. The difference is that the remainder is now small enough to name.",{"type":27,"tag":117,"props":745,"children":747},{"id":746},"the-honest-accounting-what-agentic-development-actually-costs",[748],{"type":36,"value":749},"The honest accounting: what agentic development actually costs",{"type":27,"tag":28,"props":751,"children":752},{},[753,755,760],{"type":36,"value":754},"As of mid-2026, the market pitch is: ",{"type":27,"tag":32,"props":756,"children":757},{},[758],{"type":36,"value":759},"we use AI, so — same software, half the people, half the price.",{"type":36,"value":761}," Look at what that pitch keeps and what it deletes. It keeps the compromises — usually the demo-becomes-product arc, requirements frozen at the moment of least knowledge — and deletes the humans who silently supplied the unwritten majority of every specification. And it points an optimizer at the written fraction that remains: the agent that will happily count the credit memo, at unprecedented speed. That isn't a new methodology either, as it happens — it's the oldest one, ship the throwaway to the customer, faster than it has ever been shipped.",{"type":27,"tag":28,"props":763,"children":764},{},[765],{"type":36,"value":766},"The accounting that actually holds up is less slogan-friendly. Implementation cost falls, a lot. Review cost falls further, which is the structurally interesting part, because review scarcity is what forced both of the historical compromises. Specification cost (deciding precisely what to build, and keeping the decision true as the system teaches you better) does not fall, and becomes the dominant line item. That is where the senior time goes: eliciting intent, writing it down, maintaining the harness that enforces it, and holding the verdict at every reconciliation where the spec meets the world. Total cost comes down. It does not halve just because headcount can. What a project gets that it could not get before, at any price a normal budget could pay, is the thing the shuttle group was buying: a system whose specification and implementation are the same story, kept true against every change.",{"type":27,"tag":28,"props":768,"children":769},{},[770],{"type":36,"value":771},"Because here is the ending I owe you, and it is not triumphant. The specification problem is not solved. Brooks called it the essence, and nothing above dissolves it — not the agents, not the harness, not the loop. What changed is the price of attacking it, and the price changed in exactly the direction the field's founders kept pointing: pilot systems, planned throwaways, living documents, clients confronted early with running software, every change checked against the whole. We didn't discover a new way to build software. We got the invoice down on the old one — the one that was right the whole time, the one that until now you needed a spaceship to justify — and found that the part it leaves to humans is the part that was always the actual job. Leslie Lamport (Turing Award laureate, creator of TLA+ and LaTeX) puts it this way: \"Coding is to programming what typing is to writing.\" The agents type now.",{"type":27,"tag":28,"props":773,"children":774},{},[775],{"type":36,"value":776},"Nobody was ever paying us to type. It just took fifty years and a typing machine to make that undeniable.",{"type":27,"tag":38,"props":778,"children":779},{},[],{"type":27,"tag":117,"props":781,"children":783},{"id":782},"common-questions",[784],{"type":36,"value":785},"Common Questions",{"type":27,"tag":28,"props":787,"children":788},{},[789,794,796,801],{"type":27,"tag":463,"props":790,"children":791},{},[792],{"type":36,"value":793},"Are AI coding agents a new software development methodology?",{"type":36,"value":795},"\nNo — and that's the point. The practices an agent-centric process enables were prescribed decades ago and proven on flagship programs: Royce's 1970 paper (misread as waterfall's origin) recommended a disposable pilot system and continuous customer involvement; Brooks's ",{"type":27,"tag":32,"props":797,"children":798},{},[799],{"type":36,"value":800},"Mythical Man-Month",{"type":36,"value":802}," (1975) said to plan a throwaway; \"No Silver Bullet\" (1986) prescribed prototype-driven requirements refinement; Parnas and Clements (1986) argued for rigorous living documentation and ran it on the A-7E avionics program. The space shuttle's on-board software group ran the full program at $35M a year. Agents changed the cost of that way of working, not the way.",{"type":27,"tag":28,"props":804,"children":805},{},[806,811],{"type":27,"tag":463,"props":807,"children":808},{},[809],{"type":36,"value":810},"Isn't specification-first agentic development just waterfall again?",{"type":36,"value":812},"\nNo. Waterfall's failure was never that it wrote things down — it was that nothing learned during implementation could flow back into what was written down, so requirements stayed frozen at the moment of least knowledge. Agile fixed that by making the running code the source of truth, which works until complexity outruns any human's ability to review a change against the whole system. Agents remove the dilemma: development stays iterative, but every change is reconciled against both the specification and the codebase, and the spec is updated when implementation proves it wrong. The specification is the most-edited artifact on the project, not a kickoff deliverable.",{"type":27,"tag":28,"props":814,"children":815},{},[816,821],{"type":27,"tag":463,"props":817,"children":818},{},[819],{"type":36,"value":820},"What is the difference between essential and accidental complexity?",{"type":36,"value":822},"\nFred Brooks drew the distinction in \"No Silver Bullet\" (1986). Accidental complexity is the friction of expression — translating a decision into working code, wrestling with languages, tools, and plumbing. Essential complexity is the conceptual construct itself: deciding what the rules actually are and how they interact. AI coding agents collapse the cost of the accidental part and leave the essential part where it has always been: with humans, who still must decide precisely what the system should do. On an agent-built project, that essential work becomes the dominant share of the remaining cost.",{"type":27,"tag":28,"props":824,"children":825},{},[826,831],{"type":27,"tag":463,"props":827,"children":828},{},[829],{"type":36,"value":830},"Why build prototypes that get thrown away?",{"type":36,"value":832},"\nBecause a prototype is a requirements instrument, not a first draft of the product. Running software put in front of a client in week one answers two questions a document can't: whether the builder understood the requirement, and — more important — whether the requirement itself survives being seen working. Discarding the prototype keeps demo shortcuts out of the production codebase; the historical failure of prototyping was that the throwaway shipped, because rebuilding was unaffordable. With agents, the rebuild is cheaper than the argument about keeping it.",{"type":27,"tag":28,"props":834,"children":835},{},[836,841],{"type":27,"tag":463,"props":837,"children":838},{},[839],{"type":36,"value":840},"If the test suite passes, why can an agent-built system still be wrong?",{"type":36,"value":842},"\nBecause a test suite is only the executable fraction of a specification — the part of your intent someone managed to write as automated checks. Human developers silently honored the unwritten remainder (intent, domain meaning, the obvious edge cases) because it was their own understanding. An agent optimizes exactly what is written and nothing else. The EvalPlus project demonstrated the gap empirically: grading the same model outputs against 80× more tests cut measured pass rates by as much as 19 to 29 percent and reordered model rankings.",{"type":27,"tag":28,"props":844,"children":845},{},[846,851],{"type":27,"tag":463,"props":847,"children":848},{},[849],{"type":36,"value":850},"Does test-driven development still work with AI coding agents?",{"type":36,"value":852},"\nNot the way it used to. TDD worked for two decades because the same person wrote the test and the code, silently supplying the intent the assertions never captured. An agent supplies none of it — it satisfies the written tests exactly and treats the unwritten remainder as genuinely absent. TDD with an agent is your specification with its tacit half deleted and its explicit half satisfied to the letter.",{"type":27,"tag":28,"props":854,"children":855},{},[856,861],{"type":27,"tag":463,"props":857,"children":858},{},[859],{"type":36,"value":860},"Do you still need senior engineers if agents write the code?",{"type":36,"value":862},"\nYes, and at both ends of the project. At the start, someone must elicit the goal and write it down as standards and scenarios precise enough for an optimizer to be held to. Throughout, someone must hold the verdict when the reconciliation loop flags a conflict between spec and implementation. Deciding which one is wrong is domain judgment, not pattern matching. And at the end, someone must judge which change requests fall inside the architecture's assumptions (now trivial) and which violate them (still real refactors). That judgment about what \"correct\" means was never a typing job, which is why it didn't get automated with the typing.",{"type":27,"tag":28,"props":864,"children":865},{},[866,871],{"type":27,"tag":463,"props":867,"children":868},{},[869],{"type":36,"value":870},"What happens when you come back to an agent-built system a year later?",{"type":36,"value":872},"\nIn a traditional project, the unwritten knowledge in the developers' heads starts fading the day active development stops, so a change request six or twelve months later is slower and riskier even with the same team — and you rarely get the same team. In an agent-centric project that knowledge was encoded as the work happened: requirements, standards, and test scenarios kept in sync with the code through every change. An agent picks the system up as fresh as the day work paused, so routine changes stay routine; only changes that violate foundational architecture assumptions still demand real engineering effort.",{"type":27,"tag":117,"props":874,"children":876},{"id":875},"sources",[877],{"type":36,"value":878},"Sources",{"type":27,"tag":880,"props":881,"children":882},"ul",{},[883,896,907,919,931,943,955,967,986],{"type":27,"tag":884,"props":885,"children":886},"li",{},[887,889,894],{"type":36,"value":888},"Royce, W. W. (1970). ",{"type":27,"tag":32,"props":890,"children":891},{},[892],{"type":36,"value":893},"Managing the Development of Large Software Systems.",{"type":36,"value":895}," Proc. IEEE WESCON. (The \"waterfall\" paper that warned the single-pass model \"is risky and invites failure\" and prescribed \"do it twice.\")",{"type":27,"tag":884,"props":897,"children":898},{},[899,901,905],{"type":36,"value":900},"Brooks, F. P. (1975). ",{"type":27,"tag":32,"props":902,"children":903},{},[904],{"type":36,"value":171},{"type":36,"value":906},", ch. 11, \"Plan to Throw One Away\"; and the 20th Anniversary Edition (1995), \"The Mythical Man-Month after 20 Years,\" for his partial recantation.",{"type":27,"tag":884,"props":908,"children":909},{},[910,912,917],{"type":36,"value":911},"Brooks, F. P. (1986/1987). ",{"type":27,"tag":32,"props":913,"children":914},{},[915],{"type":36,"value":916},"No Silver Bullet — Essence and Accident in Software Engineering.",{"type":36,"value":918}," IFIP / IEEE Computer.",{"type":27,"tag":884,"props":920,"children":921},{},[922,924,929],{"type":36,"value":923},"Parnas, D. L., Clements, P. C. (1986). ",{"type":27,"tag":32,"props":925,"children":926},{},[927],{"type":36,"value":928},"A Rational Design Process: How and Why to Fake It.",{"type":36,"value":930}," IEEE Transactions on Software Engineering SE-12. (Applied on the A-7E operational flight program at the Naval Research Laboratory.)",{"type":27,"tag":884,"props":932,"children":933},{},[934,936,941],{"type":36,"value":935},"Fishman, C. (1996). ",{"type":27,"tag":32,"props":937,"children":938},{},[939],{"type":36,"value":940},"They Write the Right Stuff.",{"type":36,"value":942}," Fast Company. (The on-board shuttle group: the existence proof for the full program, and its price.)",{"type":27,"tag":884,"props":944,"children":945},{},[946,948,953],{"type":36,"value":947},"Beck, K., et al. (2001). ",{"type":27,"tag":32,"props":949,"children":950},{},[951],{"type":36,"value":952},"Manifesto for Agile Software Development.",{"type":36,"value":954}," agilemanifesto.org.",{"type":27,"tag":884,"props":956,"children":957},{},[958,960,965],{"type":36,"value":959},"Naur, P. (1985). ",{"type":27,"tag":32,"props":961,"children":962},{},[963],{"type":36,"value":964},"Programming as Theory Building.",{"type":36,"value":966}," Microprocessing and Microprogramming 15.",{"type":27,"tag":884,"props":968,"children":969},{},[970,972,977,979,984],{"type":36,"value":971},"North, D. (2006). ",{"type":27,"tag":32,"props":973,"children":974},{},[975],{"type":36,"value":976},"Introducing BDD.",{"type":36,"value":978}," dannorth.net. Adzic, G. (2011). ",{"type":27,"tag":32,"props":980,"children":981},{},[982],{"type":36,"value":983},"Specification by Example.",{"type":36,"value":985}," Manning.",{"type":27,"tag":884,"props":987,"children":988},{},[989,991,996],{"type":36,"value":990},"Liu, J., Xia, C. S., Wang, Y., Zhang, L. (2023). ",{"type":27,"tag":32,"props":992,"children":993},{},[994],{"type":36,"value":995},"Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of LLMs for Code Generation.",{"type":36,"value":997}," NeurIPS 2023. arXiv:2305.01210",{"title":8,"searchDepth":999,"depth":999,"links":1000},3,[1001,1008,1016,1017,1022,1032,1033,1034],{"id":119,"depth":1002,"text":122,"children":1003},2,[1004,1005,1006,1007],{"id":126,"depth":999,"text":129},{"id":158,"depth":999,"text":161},{"id":176,"depth":999,"text":179},{"id":187,"depth":999,"text":190},{"id":220,"depth":1002,"text":223,"children":1009},[1010,1011,1012,1013,1014,1015],{"id":226,"depth":999,"text":229},{"id":251,"depth":999,"text":254},{"id":269,"depth":999,"text":272},{"id":287,"depth":999,"text":290},{"id":312,"depth":999,"text":315},{"id":330,"depth":999,"text":333},{"id":348,"depth":1002,"text":351},{"id":393,"depth":1002,"text":396,"children":1018},[1019,1020,1021],{"id":425,"depth":999,"text":428},{"id":516,"depth":999,"text":519},{"id":548,"depth":999,"text":551},{"id":594,"depth":1002,"text":597,"children":1023},[1024,1025,1026,1030,1031],{"id":618,"depth":999,"text":621},{"id":629,"depth":999,"text":632},{"id":653,"depth":999,"text":656,"children":1027},[1028],{"id":665,"depth":1029,"text":668},4,{"id":698,"depth":999,"text":701},{"id":709,"depth":999,"text":712},{"id":746,"depth":1002,"text":749},{"id":782,"depth":1002,"text":785},{"id":875,"depth":1002,"text":878},"markdown","content:ewahl:2026-08:ai-coding-agents-software-methodology.md","content","ewahl/2026-08/ai-coding-agents-software-methodology.md","ewahl/2026-08/ai-coding-agents-software-methodology","md",1789027397727]