Substrate & Intent
What kind of claim is this, and does the institution behave as if it believes it?
The Protocols by Which Honest Work Becomes False Public Reality
A diagnostic handbook for reading public artifacts: how institutional truth fails through Setup, Sale, Laundering, Trap, Accountability Failure, and Countermeasure.
The guide does not say, “It failed, therefore they lied.” It asks what the institution knew, what it published, who carried the downside, and whether correction mechanisms were allowed to bite.
What structural conditions make truth-decay likely before the claim is even made?
How was approval won?
How did internal reality become a defensible public artifact?
How was exit made politically or psychologically impossible after the original case failed?
Why does nobody pay when the claim fails?
What design rules short-circuit the lifecycle before it reaches escape velocity?
Use the triage to organize attention, not to count votes. Common symptoms in difficult projects carry less weight than rare failures of accountability geometry.
What kind of claim is this, and does the institution behave as if it believes it?
Does the physical, financial, operational, or statistical math close?
How are rates, baselines, historical investment, and audit horizons being used?
Who is being made to carry the downside, and can they fight back in time?
These are field signs, not independent votes. The recognition tests help locate the mechanism; accountability geometry and honesty markers carry the diagnosis.
Was approval won with a number too clean for the reference class, treated as ritual by insiders but measurement by outsiders?
Is a demo being sold with the strategic value of sustained operation?
Does process display replace the analysis needed to test reality?
Is the artifact persuasive at the intended reading depth and fragile below it?
Does the institution behave privately as if a different number or claim is true?
Are rates, baselines, or audit horizons hiding cumulative reality?
Has the argument shifted from future value to protecting past expenditure?
Does a new name, baseline, or personnel cycle sever the claim from consequence?
A positive diagnosis requires more than warning signs. It requires the absence, suppression, or failure of the honesty markers that should have corrected them.
The guide treats public truth-decay like a system failure. A minimal cutset may form when several conditions coexist.
This worksheet is deliberately unscored. Checking more boxes does not produce a verdict; it only records which questions you have examined.
The appendices demonstrate the method without pretending to statistically validate it. Their evidentiary classes are intentionally different.
Closed, high-observability infrastructure case testing cost, schedule, and accountability failure while preserving the distinction between technical success and public-consent failure.
Long-horizon cumulative case testing category fragmentation, debt service, veterans’ obligations, and audit horizons longer than political horizons.
Shows how the same tools behave in software/AI and macroeconomic communication without treating either caselet as a full empirical conviction.
A strategically distorted Sale followed by accountability geometry strong enough to arrest the sequence and recover the public record.
When root cause and investigation evidence are controlled by the subject, the guide returns an interval rather than a point verdict.
Separates causal adjudication from institutional-conduct review so over-reassurance is not confused with proof of pesticide causation.
The web presentation above is a reading layer. The complete source text follows below, including its caveats, case studies, tables, research agenda, and bibliography.
This review build reorganizes the presentation for the web while preserving the v1.0 source framework and its terminology. It does not silently convert hypotheses into measurements, collapse stage-specific verdicts into program-level convictions, or remove the guide’s false-positive and observability-limit guardrails.
A FIELD GUIDE TO INSTITUTIONAL DISHONESTY
The Protocols by Which Honest Work Becomes False Public Reality
v1.0 - Clean Release with Appendix F Boundary Case
A diagnostic handbook for reading public artifacts: how institutional truth fails through Setup, Sale, Laundering, Trap, Accountability Failure, and Countermeasure.
Institutional dishonesty is what happens when authority, consequence, language, and publication rights are arranged so honest analysis cannot survive contact with approval. That is the thesis of the field guide. The problem is not merely bad people lying. The problem is institutional architecture that can transform honest internal work into a misleading public artifact while allowing most participants to experience themselves as responsible professionals.
The central claim: institutional dishonesty exists when a public artifact materially diverges from what the institution had reason to know, and when that divergence helps secure approval, funding, legitimacy, or delay while externalizing cost onto people outside the decision room.
This guide is not a claim that every failed estimate is a lie. Complex work contains real uncertainty. The sharper claim is that certain institutional conditions predictably convert uncertainty into selective public certainty: optimistic baselines, milestone inflation, filtering of internal warnings, ceremonial governance, renaming, rebaselining, and sunk-cost entrapment.
The framework is organized as a lifecycle. First the institution sets up conditions for truth-decay. Then it sells the claim. Then it launders the public artifact. Then it traps the funder after the original case degrades. Finally, it escapes accountability through time, rotation, rebaselining, or fragmentation. The named specimens are not treated as perfectly disjoint species; they are eight primary failure modes with several named sub-variants, because real institutional behavior is mechanically entangled.
The guide is deliberately stage-specific. The specimens are warning vocabulary, not votes. Accountability geometry, honesty markers, ordinal specificity, and observability limits carry the diagnostic burden; numeric base-rate tables remain future work, not a claim smuggled into the method.
The intellectual roots are Flyvbjerg on megaproject overruns, Meyer and Rowan on institutional ceremony, Brunsson on organizational hypocrisy, Jackall on bureaucratic moral selection, Frankfurt on bullshit, public-choice and principal-agent theory on incentives, and Perrow on system accidents. The distinctive move here is to treat institutional truth-decay as a PRA-style fault tree: a public artifact fails when consequence diffusion, delayed auditability, filtered warnings, and legitimacy pressure form a minimal cutset.
The appendices demonstrate the method without pretending to validate it statistically. Vogtle and war finance are positive applications. JWST is a red-team case: a distorted Sale was arrested by real accountability geometry. New Glenn is an open-verdict case: when the evidence is controlled by the subject and root cause is unknown, the honest output is an interval, not a point verdict.
This guide is meant for reading institutional artifacts: budget requests, regulatory filings, strategic plans, technology roadmaps, cost estimates, public statistics, milestone announcements, program reviews, war authorizations, and after-action reports.
Use it as one procedure with three nested levels.
Do not infer dishonesty from failure alone. A bridge can run late because geology surprised everyone. A war can cost more because the adversary adapts. A software pilot can fail because the problem was harder than it looked. The field guide becomes relevant when the public claim diverged from what reasonable internal evidence, historical reference classes, or institutional behavior already suggested at the time approval was sought.
A failed estimate becomes institutionally dishonest only when the public claim materially diverged from information reasonably available to the institution at the time of approval, and when that divergence benefited the institution while externalizing cost, risk, or delay onto others.
This rule prevents the framework from becoming a conspiracy machine. The guide does not say, “It failed, therefore they lied.” It asks whether the institution had reason to know the claim was presented on the wrong substrate, whether the uncertainty was disclosed before consent, whether the internal behavior contradicted the public statement, and whether the people paying the price had any meaningful recourse.
The guide should not be used as a checklist where every warning sign counts as one vote for dishonesty. That is the common-cause error: several specimens can light up because the same large, novel, difficult project generated all of them. Counting correlated symptoms as independent evidence overstates the diagnosis.
The better procedure is ordinal specificity reasoning. Numeric likelihood-ratio language is useful as an intuition, but this guide does not yet have the base-rate data to compute real ratios. It therefore ranks observations qualitatively as low-, medium-, or high-specificity pending future research. Common overruns, rebaselines, schedule slips, and technical novelty move the diagnosis only slightly. Suppressed review, hidden baselines, refusal to publish unit costs, fake off-ramps, and vanished original baselines move it more.
In practice: treat the specimens as field signs, not verdicts. Treat honesty markers and accountability geometry as the load-bearing test, and return a verdict by lifecycle stage rather than by affection or dislike for the whole program.
Low-specificity observations include ordinary overrun, rebaseline, schedule slip, technical novelty, and optimistic early estimates. High-specificity observations include filtered reviews, hidden baselines, refusal to publish unit costs, non-binding caps, fake off-ramps, disappearing original baselines, and behavior/public-claim divergence that persists after correction was possible. The guide’s current tool is ordinal weighting, not arithmetic. Building numeric base rates by habitat is a research deliverable, not a premise already proven.
The evidence is part of the system being judged. In many important cases, the institution controls the records, logs, root-cause narrative, dissent channels, and public archive by which outsiders would judge it. A dishonest system can therefore produce not only bad outcomes but bad records of those outcomes. This is a common-cause failure between the system and its instrumentation: the same error as treating correlated sensor readings as independent confirmation.
The key inversion: observability often falls as the thing being hunted rises. A competent, well-resourced institution with strong narrative control may leave less visible evidence than a clumsy one. The absence of damning evidence - no whistleblower, leaked memo, or recorded dissent - is therefore near-zero-specificity by itself. It may mean nothing was wrong, that evidence was buried, or that the organization no longer produces dissent.
A clean audit trail is similarly limited. It can show that a tool was used as authorized; it does not prove the authorization was legitimate. A perfectly logged system can execute an illegitimate policy flawlessly. The log documents conduct within the frame, not the legitimacy of the frame.
The disciplined response is to interrogate the architecture that would create or suppress evidence: independent review, publication rights, barrier scope, dissent channels, binding caps, unit-cost visibility, and whether external review actually engaged. These structural facts are harder for the subject to launder than its narrative about what happened.
Self-check: Am I reading a verdict from an absence? Have I confused “nothing was found” with “nothing is there”? Is my confidence narrower than the subject’s control over the record permits?
Field test: when narrative control is high, return an interval, not a point verdict. The width of the interval is itself a finding.
False-positive discipline prevents the guide from convicting every failure. False-negative discipline prevents it from acquitting every success. Ordinal specificity prevents it from counting correlated symptoms as independent votes. Observability-limit discipline prevents the subtlest error of the four: mistaking the limits of what the institution allowed the analyst to see for the limits of what actually occurred.
Front Matter - Observability Limit Discipline: why evidence control limits verdict confidence.
Part I - Foundations: the setup conditions for institutional truth-decay.
Part II - Lifecycle Map: how the mechanisms fit together.
Part III - Specimens: the named mechanisms and fair caveats.
Part IV - Diagnostics: the field checklist and triage flow.
Part V - Habitats: where the protocol thrives and where it is weaker.
Part VI - Responses: what honest analysts and institutions can do.
Part VII - Case Studies: placement, case-study template, and appendices.
Appendix D - Red-Team Case Study: James Webb Space Telescope.
Appendix E - Live-Case Study: Blue Origin New Glenn Static-Fire Anomaly.
Appendix F - Two-Question Boundary Case: Zika, pyriproxyfen, and institutional over-reassurance.
Part VIII - Research Agenda and Bibliography.
The guide must pass its own tests. Several of its strongest claims are diagnostic hypotheses, not settled measurement claims. For example, consequence diffusion is a strong candidate predictor of institutional dishonesty, but it is not yet a quantified law. It also has a confound: diffusion often correlates with scale, and scale itself predicts overrun. Therefore consequence diffusion may predict difficulty, overrun, or weak correction before it specifically predicts dishonesty. The framework must not dress this hypothesis in measurement-substrate clothing before the data support it.
This document therefore makes three levels of claim. First, some claims are empirical and source-bound, such as cost and schedule outcomes in the Vogtle, war-finance, and JWST cases. Second, some claims are analytic interpretations of public artifacts. Third, some claims are research hypotheses that should be tested against a broader case matrix. The guide is strongest when it labels which kind of claim it is making.
The analyst is also part of the system. A determined reader can find optimism, diffuse consequence, ceremonial process, and time drift in almost any complex program. Those features are common in innocent difficulty as well as dishonest architecture. The field guide becomes dangerous if the vivid specimen names are kept while the burden-of-proof rule and honesty markers are skipped.
Misuse warning: the guide must not become a worldview installation device. It is easy to pick a target, harvest the specimen vocabulary, and skip the exonerating test. It is also easy to protect a loved program by over-weighting its final success and skipping the front-end sale. The required discipline is bidirectional: start with base-rate awareness, search for both dishonesty markers and honesty markers, and ask what evidence would move the case away from or toward dishonesty before looking for a conclusion.
A diagnosis should be downgraded when the institution published ranges before approval, preserved the original baseline, disclosed credible reference classes, named risk owners, separated demonstration from operation, updated the public record without rebranding, and allowed adversarial review to publish dissent. Those behaviors do not guarantee honesty, but they are evidence against the protocol.
Self-check: Am I treating failure as proof, or am I showing divergence between the public claim and information reasonably available at the time? Am I grading a beloved successful program on the curve of its outcome? Am I using base-rate awareness and ordinal specificity, or just pattern-matching? Am I distinguishing ordinary uncertainty from strategic misrepresentation stage by stage? If not, the guide is being misused.
Honesty markers: before elevating a diagnosis, look for evidence that the institution allowed reality to bite. Independent reviews with publication rights, real off-ramps, binding caps, reauthorization after breach, public rebaselines that preserve the original baseline, behavioral convergence after correction, and personnel continuity across the accountability window all downgrade the finding. They do not erase failure, but they can move the case from institutional dishonesty to ordinary uncertainty or optimism bias.
Field test: Am I asking whether the program as a whole is dishonest, or am I asking which lifecycle stage behaved dishonestly and which stage corrected it?
The analyst must guard both sides: do not convict every failure, and do not acquit a beloved success merely because it eventually delivered. Diagnose the lifecycle stage, not the emotional valence of the final outcome.
The JWST red-team case forces a hierarchy and a temporal discipline. Surface specimens identify places to look. Honesty markers decide whether the institution allowed reality to bite. A case can show optimistic baselines, rebaselines, sunk-cost pressure, and reading-depth defense and still remain below institutional dishonesty in later stages if the correction mechanisms stayed live and were actually exercised. But that later accountability success does not automatically cleanse the original sale.
The six primary honesty markers are:
A positive diagnosis requires more than the presence of warning signs. It requires the absence, suppression, or failure of the honesty markers that should have corrected them.
There are two fundamentally different kinds of numbers. A measurement-substrate number describes a reality that is indifferent to the claim being made about it. The speed of light does not change because a committee announces a more convenient value. A bridge either carries the load or it does not. A reactor either closes its heat balance or it does not.
A faith-and-effort number describes an outcome partly created by belief, commitment, and coordination. A startup revenue target, a fundraising goal, a sales forecast, or a wartime mobilization target can become more achievable because people organize around it. In this substrate, confidence is not merely descriptive. It is causal.
Both substrates are legitimate. The fraud-like move begins when a faith-and-effort number is dressed in measurement language and used to obtain approval from people who are not told which substrate they are standing on.
This connects to Harry Frankfurt’s analysis of bullshit: the danger is not always a direct lie, but speech detached from adequate regard for truth. In institutional settings, the output may be truth-indifferent while still sounding technical, reviewed, and responsible.
Field test: Is this number being presented as measurement, aspiration, negotiation, morale, or political positioning? Would the audience interpret it differently if the substrate were stated plainly?
The best first place to look in this framework is the separation between the person making the commitment and the person who pays when the commitment fails. It is a screening heuristic, not a fully quantified predictor. Consequence diffusion repeatedly appears at the root of the case studies and in adjacent literatures: public choice, principal-agent theory, and collective-action analysis. But it also confounds with scale, and scale itself predicts difficulty and overrun; see “Applying the Guide to Itself.”
This overlaps with public choice theory, Olson’s logic of concentrated benefits and diffuse costs, and principal-agent theory. The public may be the principal in theory, but if it cannot observe, exit, sue, refuse payment, or discipline the agent in time, its nominal authority is mostly decorative.
Danger pattern: concentrated authority, concentrated reputational benefit, diffuse financial consequence, delayed auditability.
Field test: Who pays if this number is wrong? Were they represented when it was approved? Can they discipline the claim before the damage is locked in?
Horizon Drift occurs when the time required to verify a claim exceeds the time during which the approving actors can still be held to it. In that design, the system is insulated from reality at birth. The project, war, pension promise, infrastructure plan, or technology roadmap may not become fully auditable until after the decision seats have rotated and the original approval narrative has hardened into institutional memory.
This is why Horizon Drift belongs in both Setup and Accountability Failure. It is a setup condition when delay is built into the claim from the beginning. It is an accountability failure when delayed audit, personnel rotation, archive loss, or rebaselining prevents the original decision from being judged against the original public promise.
Field test: Will the claim become auditable while the approving seats can still be named, reached, and disciplined? If not, the system is born with accountability drift.
Institutional theory adds a third foundation. Meyer and Rowan showed how formal structures can operate as myths and ceremonies: they provide legitimacy while decoupling from actual work. A document can therefore be procedurally impressive and operationally hollow.
The field guide’s concern is closure. Does the artifact contain the balances necessary to test reality: mass, energy, cost, schedule, maintenance, staffing, financing, and accountability? Or does it display governance in place of proof?
Field test: Does this artifact prove capability, or does it display legitimacy?
Perrow’s system-accident framing matters because institutional dishonesty does not always begin with a villain. In complex, tightly coupled systems, bad outcomes can emerge from normal operation. The same is true for truth. Legal review, executive caution, budget positioning, public affairs discipline, and partner coordination may each be locally rational. Together they can produce a public artifact nobody inside would rely on privately.
The institution may contain honest people, honest models, and honest warnings. The question is whether those things survive publication.
The framework can be read as a Probabilistic Risk Assessment applied to public truth. The top event is not core damage or system loss; it is a materially misleading public artifact that survives long enough to govern real money, authority, or public consent.
The lifecycle stages are the failure sequence. The specimens are failure modes. The diagnostics are closure tests. The countermeasures are barriers. A minimal cutset might look like this: consequence diffusion AND delayed auditability AND filtered warnings AND approval pressure. When those conditions coexist, the system can produce dishonesty without requiring every participant to intend dishonesty.
This PRA framing is useful because it prevents moral laziness in both directions. It does not excuse the outcome as accidental, but it also does not require a cartoon villain. It asks which barriers failed, which warnings were filtered, which incentives aligned, and which cutset allowed the public artifact to detach from reality.
This map organizes the mechanisms by lifecycle stage rather than treating them as equal fragments. That makes the framework usable on a live artifact: first locate the stage, then test whether the relevant barriers fired or failed.
| Lifecycle Stage | Primary Mechanisms | Core Question |
|---|---|---|
| Setup | Measurement/Faith-and-Effort Split; Consequence Diffusion; Horizon Drift; Legitimacy vs. Closure | What structural conditions make truth-decay likely before the claim is even made? |
| The Sale | Optimistic Baseline & Handshake Protocol; Demonstrate/Operate/Sustain Conflation | How was approval won? |
| The Laundering | Legitimacy Laundering; Reading-Depth Defense; Specificity Inversion | How did internal reality become a defensible public artifact? |
| The Trap | Sunk-Cost Colonization; Uncertainty Scapegoat; Integral/Derivative Framing | How was exit made politically or psychologically impossible after the original case failed? |
| Accountability Failure | Renaming and Accountability Escape; Horizon Drift; Archive Failure; Personnel Rotation | Why does nobody pay when the claim fails? |
| Countermeasure | Reference Classes; Publication Rights; Mandatory Re-Gating; Seat/Title Baseline Tracking | What design rules short-circuit the lifecycle before it reaches escape velocity? |
Horizon Drift appears twice in this lifecycle on purpose. It can be a birth defect in the setup when auditability is structurally delayed from the start, and it can be an accountability failure when personnel rotation, rebaselining, or delayed cost recognition prevents the original decision from being judged against its original claim.
The specimens are not independent votes, and they should not be summed. They are overlapping views of a smaller mechanism set, like fault-tree branches that can share common causes. Optimistic Baseline, Sunk-Cost Colonization, Rebaselining, and Reading-Depth Defense may all appear because the same underlying project is large, novel, late, and politically protected. The diagnostic question is not “how many specimens fired?” but “which observations have enough specificity to move the case after considering the domain’s ordinary base-rate problems?”
Use the specimens as a vocabulary for describing the failure sequence. Use accountability geometry, high-specificity observations, and the burden-of-proof rule to classify the case.
The approval case is built around a number that is too low, too fast, or too clean relative to the relevant reference class. The Handshake Protocol is the ritual version: requester and approver both understand the number is not the real expectation, but it is the formal language required to begin. Flyvbjerg’s megaproject literature is the load-bearing reference here: his work repeatedly shows large projects tending to go over budget and over schedule, and argues that optimism bias and strategic misrepresentation are central causes.
Field test: Was the reference class applied before approval, or only discovered after the project became irreversible? Fair caveat: a baseline can be wrong without being dishonest; the issue is knowable contrary evidence excluded from the approval case.
A single milestone collapses three different capability tiers: demonstrate a function, operate it reliably, and sustain it as a net-positive part of the larger system. The budget buys the demo while the strategic justification borrows the value of sustained operation.
Field test: Which tier is funded, which tier is tested, and which tier is used to justify the program? Fair caveat: demonstrations are legitimate; the deception begins when demonstration language implies operational maturity.
This combines the earlier Farm Filter and Ceremonial Governance mechanisms. Honest internal analysis passes through layers of legal, public-affairs, executive, partner, budget, and reputational review until the public artifact becomes procedurally defensible but substantively misleading. Governance display is part of the laundering: review boards, acronyms, dashboards, and compliance rituals create legitimacy while avoiding closure math.
Field test: Does the artifact publish the analysis that would let an outsider test the claim, or does it mainly display process? Fair caveat: some details are legitimately withheld; laundering exists when omission prevents evaluation of feasibility, cost, or consent.
A document is engineered to be persuasive at the depth it expects to be read: too shallow and it looks impressive, too deep and the closure gaps appear. Specificity Inversion is the build technique: the artifact is precise where specificity is cheap and vague where specificity would be expensive.
Field test: Is the document specific about vendors, milestones, dashboards, and governance, but vague about cost basis, maintenance burden, risk owners, or physical closure? Fair caveat: not every missing detail is evasion; the pattern matters.
The institution behaves privately as if one claim is true while publishing another. Quadruplespeak is the numeric/audience-segmentation variant: stated cost, acknowledged cost, actual cost, and true social or lifecycle cost. The term is deliberately vivid, but the sober mechanism matters more than the label: different audiences receive different cost realities while the institution preserves procedural defensibility.
Field test: What number do internal planners behave as if they believe, and does that match the number presented to the funding public? Fair caveat: internal contingency planning is not itself deception; deception begins when the contingency is the real expectation and the public artifact hides it.
The institution manipulates time, baselines, and auditability. Integral vs. Derivative Framing emphasizes the rate while the public experiences the cumulative level. Horizon Drift separates the time of approval from the time of audit. The result is a claim that can be true in a narrow current-frame sense while false in the lived cumulative sense.
Field test: Is the headline a rate while the consequence is cumulative? Will the claim become auditable before responsible seats rotate? Fair caveat: rates and long horizons can be technically appropriate; the issue is their use to hide accumulated consequence.
After the original approval case fails, the argument shifts from “this is worth doing” to “we cannot stop now.” The Optimistic Baseline is the bait; Sunk-Cost Colonization is the trap. The Uncertainty Scapegoat is the defense variant: uncertainty that should have expanded disclosure before consent is invoked after failure to protect continuation.
Field test: Has the decision standard changed from forward-looking value to preservation of past expenditure? Fair caveat: continuation can be rational after sunk cost; the error is treating past spending as proof of future value.
A continuing activity is relabeled so the accountability clock resets. Personnel rotation, archive failure, rebaselining, and program reclassification all perform the same escape function: the old claim disappears while the underlying activity continues. In software and AI programs, Agile, MVP, and pivot language can become the socially acceptable vocabulary of this escape when used to avoid a hard closure test.
Field test: What accountability mechanism would apply if the old name, original baseline, and original decision seats remained visible? Fair caveat: programs sometimes genuinely change; the warning sign is continuity of substance paired with discontinuity of accountability.
| Stage | Primary Mechanism | One-Line Recognition Test |
|---|---|---|
| Sale | Optimistic Baseline & Handshake Protocol | Was approval won with a number too clean for the reference class, treated as ritual by insiders but measurement by outsiders? |
| Sale | Demonstrate / Operate / Sustain | Is a demo being sold with the strategic value of sustained operation? |
| Laundering | Legitimacy Laundering | Does process display replace the analysis needed to test reality? |
| Laundering | Reading-Depth Defense | Is the artifact persuasive at the intended reading depth and fragile below it? |
| Laundering | Public/Private Claim Split | Does the institution behave privately as if a different number or claim is true? |
| Trap | Temporal Framing | Are rates, baselines, or audit horizons hiding cumulative reality? |
| Trap | Sunk-Cost Colonization | Has the argument shifted from future value to protecting past expenditure? |
| Accountability Failure | Renaming and Accountability Escape | Does a new name, baseline, or personnel cycle sever the claim from consequence? |
The same lifecycle behaves differently in different physical regimes. Software and AI programs often deceive by treating a scripted, sandboxed prototype as if it were scalable operating infrastructure. Heavy infrastructure often deceives by treating a massive, uncertain, decades-long industrial build as if it can be locked into a clean baseline. In software, the vocabulary of Agile, pivoting, MVP iteration, and product discovery can become a built-in Renaming Gambit: each failed closure test is reframed as learning, iteration, or a new product layer. In heavy atoms, the physical plant cannot be renamed into existence; the lie usually lives in cost, schedule, financing, and permitting assumptions.
| Domain | Typical Demonstration Shell Game | Typical Failure Mode |
|---|---|---|
| Software / AI | A sandboxed prototype, scripted demo, or “MVP” is presented as evidence of scalable, maintainable operation. | Integration burden, data drift, hallucination, human fallback, token cost, governance debt, security review, and Agile/pivot vocabulary used to defer the hard closure test. |
| Heavy infrastructure | A preliminary cost/schedule baseline is presented as if permitting, supply chain, fabrication, labor, financing, and rework will behave. | Schedule drift, quality problems, financing cost, contractor failure, and regulatory rework consume the business case. |
The diagnostics are grouped as a chronological triage flow, but they must be read through domain base-rate awareness, ordinal specificity, and honesty markers. A common warning sign in a high-novelty domain has less diagnostic weight than a rare accountability failure. Use the checklist to organize attention, not to count votes.
| Triage Stage | Tests | Question |
|---|---|---|
| 1. Substrate & Intent | Substrate Test; Internal-Belief Test; Public/Private Split | What kind of claim is this, and does the institution behave as if it believes it? |
| 2. Engineering Ledger | Balance-Sheet Test; Tier Test; Specificity-Inversion Test; Reference-Class Test | Does the physical, financial, operational, or statistical math close? |
| 3. Temporal Shell Game | Integral Test; Clock Test; Horizon Test; Sunk-Cost Test; Baseline Test | How are rates, baselines, historical investment, and audit horizons being used to hide cumulative drift or delay judgment? |
| 4. Extraction Profile | Consequence Test; Recourse Test | Who is being made to carry the downside, and can they fight back in time? |
The earlier “three passes” and four-stage triage are the same procedure at different zoom levels. First, scan the artifact. Second, run the triage. Third, place the finding in the lifecycle.
1. Artifact scan: identify the claim, audience, substrate, consequence path, and audit horizon. 2. Triage flow: test Substrate & Intent, Engineering Ledger, Temporal Shell Game, and Extraction Profile. 3. Lifecycle placement: decide whether the claim is in Setup, Sale, Laundering, Trap, Accountability Failure, or Countermeasure.
Step 0 - Base-rate awareness, honesty markers, and observability boundaries: before scoring warning signs, ask how common the apparent failure mode is in this domain, whether the institution preserved the mechanisms that could correct it, and whether the subject controls the sensors and records by which it would be judged. Do not pretend to compute numeric likelihood ratios unless the relevant base-rate data actually exists.
[ ] Observability boundaries: Does the subject institution control the sensors, logs, root-cause record, or dissent channel for the event? If yes, widen the confidence interval and test the data-producing architecture before scoring the case.
[ ] Substrate: Is the claim measurement, aspiration, negotiation, morale, or positioning?
[ ] Consequence: Who pays if it is wrong, and were they in the room?
[ ] Closure: Does the artifact publish the relevant mass, energy, cost, schedule, maintenance, and accountability balances?
[ ] Tier: Does it distinguish demonstrate, operate, and sustain?
[ ] Specificity: Is precision concentrated where it is cheap and absent where it matters?
[ ] Reference class: What happened in comparable projects, and was that history used before approval?
[ ] Behavior: What does the institution budget, staff, schedule, and contingency-plan as if it believes?
[ ] Integral: Is the headline a rate while the lived consequence is cumulative?
[ ] Clock: Can relabeling, rebaselining, or reauthorization reset accountability?
[ ] Sunk cost: Has the argument shifted from future value to past expenditure?
[ ] Horizon: Will the claim become auditable before responsible people leave?
[ ] Recourse: Can the affected population exit, refuse, sue, vote effectively, or discipline the decision before the cost is locked in?
[ ] Honesty markers: Did independent review, publication rights, off-ramps, binding caps, reauthorization, personnel continuity, or behavioral convergence remain live and exercised?
The protocol is strongest where outcomes are sensitive to cost or cumulative burden, consequences are diffused onto people outside the decision room, and independent verification is weak, delayed, or too technical for the affected public to use.
| Habitat | Why Vulnerable | Common Specimens |
|---|---|---|
| Government procurement and infrastructure | Diffuse funders, long timelines, weak recourse, political approval thresholds, and delayed auditability. | Optimistic Baseline & Handshake; Legitimacy Laundering; Public/Private Claim Split; Sunk-Cost Colonization. |
| Nuclear construction | Capital discipline determines economics; schedule drift and financing cost can destroy the public-consent case even when the physics works. | Optimistic Baseline & Handshake; Temporal Framing; Sunk-Cost Colonization; Renaming and Accountability Escape. |
| War finance | The full bill arrives decades after the authorization moment; categories fragment across agencies, debt service, veterans’ care, and social cost. | Temporal Framing; Public/Private Claim Split; Horizon Drift; Renaming and Accountability Escape. |
| Macroeconomic communication | Technical indices can be serious while public rhetoric overuses them as complete proxies for lived economic experience. | Temporal Framing; Reading-Depth Defense; Public/Private Claim Split. |
| Prestige technology programs | Demos are cheap, integration is expensive, executive rewards arrive before TCO, and Agile/MVP language can reset the accountability clock. | Demonstrate/Operate/Sustain; Renaming and Accountability Escape; Legitimacy Laundering; Public/Private Claim Split. |
Fundamental physics and mathematics are more resistant because reality and proof cannot be lobbied in the same way a budget committee can. But this immunity has boundaries. The core data may remain measurement-substrate while the funding interface shifts into faith-and-effort language. Big science can preserve clean physics while filtering the cost, schedule, and political case used to obtain billions of dollars.
The same boundary matters in nuclear power. The nuclear physics did not bend at Vogtle. The financial interface did. A machine can split atoms successfully while the institutional framework around its construction vaporizes public trust.
No private analyst can single-handedly reform a large institution. But honest analysis can still make the protocol more expensive, more visible, and harder to launder.
Produce the analysis the institution cannot publish: realistic ranges, closure math, reference-class comparison, consequence maps, and original-baseline comparisons.
Refuse euphemism when the burden-of-proof rule is met. If the number was known to be wrong when used for approval, call the public artifact dishonest.
Archive contemporaneous warnings, but do not stop at abstract entities. Tie estimates, approvals, warnings, and reversals to job titles, institutional seats, committee roles, and decision records. Modern institutions can live comfortably in documented irony unless the archive preserves accountability geometry.
Each capable analyst who refuses to perform the protocol withdraws cognitive capacity from it. The institution may recruit lesser validators, but that degradation matters.
Countermeasures are strongest when mapped to the lifecycle stage they interrupt. A generic reform list is easy for an institution to absorb. A stage-specific trigger is harder to evade because it asks the right question at the moment when the protocol needs silence.
| Stage | Trigger | Blocks |
|---|---|---|
| Setup | Consequence map; named seats; audit-horizon test. | Birth-defect claims whose audit arrives after responsibility expires. |
| Sale | Reference-class forecast; published range; independent estimate. | Optimistic Baseline and Handshake Protocol. |
| Laundering | Red-team publication rights; preserved dissenting appendices. | Legitimacy Laundering and Reading-Depth Defense. |
| Trap | Mandatory sunk-cost re-gating using forward-looking value only. | Sunk-Cost Colonization after major drift. |
| Accountability | Original-baseline tracking by seat/title; outcome postmortem. | Personnel rotation, Horizon Drift, archive failure, and Renaming Gambit. |
| Public recourse | Plain-language consequence statement; cumulative burden report. | Diffuse cost before it becomes irreversible. |
The full case studies live in the appendices. They are not presented as a representative validation sample. Vogtle and war finance are strong positive applications; JWST is a demanding negative control; New Glenn is an open-verdict observability test; pyriproxyfen is a two-question boundary case separating causal adjudication from institutional-conduct review; the AI and CPI caselets are transferability sketches. The honest evidentiary claim at this stage is that the tool can be applied coherently across habitats, not that it has been statistically validated against a representative case universe.
Future validation requires a case matrix selected before diagnosis, with positive cases, negative controls, and ambiguous cases across comparable domains. Until then, the appendices are demonstrations of method, not proof of universal frequency.
The appendices are different evidentiary classes, not a uniform sample.
| Appendix | Case Class | Methodological Use |
|---|---|---|
| Appendix A - Vogtle 3 & 4 | Positive application (closed case, high observability) | Tests cost/schedule/accountability failure in heavy infrastructure where public outcome data is available. |
| Appendix B - Post-9/11 War Finance | Positive application (long-horizon cumulative case) | Tests boundary definition, horizon drift, category separation, and cumulative cost accounting. |
| Appendix C - AI and CPI Caselets | Transferability sketches (interpretive applications) | Shows portability to software/AI and macroeconomic communication without treating them as full empirical convictions. |
| Appendix D - JWST | Negative control (closed case, verified accountability geometry) | Shows that a difficult, overrun program can have a strategically distorted Sale while later accountability barriers fire and prevent permanent public-record distortion. |
| Appendix E - New Glenn | Open-verdict boundary test (live case, low observability) | Shows how the guide should behave when the root cause, dissent channel, and investigation record are controlled by the subject and the honest output is an interval, not a point verdict. |
Appendix A is a positive application: a closed, high-observability infrastructure case.
Appendix B is a positive application: a long-horizon cumulative-finance case.
Appendix C is a transferability sketch: software/AI and CPI caselets that test habitat differences.
Appendix D is a negative-control and red-team case: a strategically misrepresented Sale whose accountability barrier fired.
Appendix E is an open-verdict boundary test: live case, low observability, interval output.
Appendix F is a two-question boundary case: the causal question resolves mostly toward Zika, while the institutional-conduct question tests whether public reassurance outran a triggered toxicology burden.
| Section | Purpose |
|---|---|
| Claim at Approval | What was publicly sold? |
| Known or Knowable Baseline | What historical or internal information should have disciplined the claim? |
| Lifecycle Map | Which mechanisms appeared during sale, laundering, trap, and accountability escape? |
| Consequence Map | Who gained from the claim, who paid for the miss, and what recourse existed? |
| Falsification Test | What facts would have made this ordinary uncertainty rather than institutional dishonesty? |
| Lessons for the Guide | What does the case prove or pressure-test? |
Plant Vogtle Units 3 and 4 are a clean case because the result is mixed rather than cartoonish. The reactors ultimately entered commercial operation and provide large-scale, low-carbon electricity. In physical and climate terms, that matters. But the project also became a cost, schedule, and public-consent failure: the public approval case did not resemble the final burden carried by ratepayers and stakeholders.
This distinction is essential. The case is not anti-nuclear. It is anti-decoupling. A reactor can be technically successful while the institutional truth-system around its construction fails.
| Dimension | Public / Early Case | Later Reality |
|---|---|---|
| Cost | Early reporting and approval-era expectations centered around roughly $14 billion for the expansion. | Later public estimates exceeded $30 billion, depending on cost boundary and accounting treatment. |
| Schedule | Commercial operation was originally expected in the 2016-2017 range. | Unit 3 entered commercial operation in July 2023; Unit 4 entered commercial operation in April 2024. |
| Technology | AP1000 deployment was treated as a path toward a new standardized nuclear build era. | The project exposed severe first-of-a-kind, supply-chain, contractor, quality, regulatory, and financing risks. |
| Public burden | The project was sold as an asset that would justify its cost. | Ratepayer recovery became a central political and regulatory issue. |
| Lifecycle Stage | Vogtle Pattern |
|---|---|
| The Sale | Optimistic Baseline and Handshake Protocol: the project required an approval-era cost/schedule case that did not reflect the actual execution burden of first-of-a-kind nuclear construction in the United States. |
| The Laundering | Ceremonial Governance and Reading-Depth Defense: the project was review-heavy and procedurally elaborate, but the public-facing confidence did not force ordinary ratepayers to confront the full capital-discipline risk. |
| The Trap | Sunk-Cost Colonization: after major expenditure and the Westinghouse bankruptcy, the argument naturally shifted from “this is cheap enough to build” to “stopping now would waste what has already been spent.” |
| Accountability Failure | Consequence Diffusion and Horizon Drift: the approval benefits were concentrated among institutional actors and policy goals, while the financial consequences diffused across ratepayers and long time horizons. |
Vogtle shows why Sunk-Cost Colonization deserves to be a named dynamic. The Optimistic Baseline is the bait. Sunk-Cost Colonization is the trap. Once billions have been spent, the project no longer has to win on the original terms. It only has to make cancellation feel worse than continuation.
That does not mean continuation was irrational. A serious case can exist for finishing a partially built reactor. The point is narrower: after enough money is committed, the public is no longer evaluating the same decision it was asked to approve. The decision has been colonized by prior expenditure.
A critic should be invited to falsify the institutional-dishonesty reading. Vogtle would look more like ordinary uncertainty if the approval case had prominently disclosed reference-class nuclear construction risk, preserved the original baseline in public reporting, clearly allocated downside burden, separated technical feasibility from capital-discipline risk, and forced periodic fresh go/no-go decisions not dominated by sunk-cost logic.
The challenge for defenders is not to show that nuclear energy is valuable. It is. The challenge is to show that the public-consent artifact honestly represented the cost, schedule, and burden structure at the time decisions became difficult to reverse.
War finance is the archetype of accountability horizon failure. The political decision is made in the present, but the full cost becomes visible over decades: operations, reconstruction, equipment replacement, base-budget growth, debt service, veterans’ care, disability payments, and opportunity cost. By the time the integral is auditable, the original decision-makers, authorizers, and public narratives have moved on.
As of the Costs of War findings page consulted for this version
| Cost Boundary | What It Captures | Field-Guide Issue |
|---|---|---|
| Near-term appropriations | Annual operations, deployments, equipment, and agency spending. | Derivative framing: what is spent this year can hide the cumulative bill. |
| Budgetary lifecycle cost | Past appropriations plus future obligations such as veterans’ care and disability. | Integral reality: the burden continues after the political moment. |
| Debt service | Interest costs when war spending is deficit-financed. | Quadruplespeak: financing cost is real but often excluded from headline war cost. |
| Social and opportunity cost | Lives, injuries, displacement, family burden, foregone domestic investment. | Consequence diffusion: much of the burden is not carried by the decision-makers. |
| Lifecycle Stage | War-Finance Pattern |
|---|---|
| The Sale | Initial debate emphasizes mission, necessity, urgency, and near-term funding. Full lifecycle costs are uncertain, but uncertainty tends to reduce specificity rather than expand public cost ranges. |
| The Laundering | Cost categories fragment across supplemental appropriations, base budgets, veterans’ care, debt service, intelligence, reconstruction, and future obligations. The whole cost exists, but not in one politically salient number. |
| The Trap | Sunk-Cost Colonization appears as the argument shifts from “this will be limited” to “we cannot leave because prior sacrifice would be wasted.” |
| Accountability Failure | The audit horizon exceeds the political horizon. The public learns the integral after the offices, committees, narratives, and authorizations have changed. |
War finance shows that the field guide is not merely a construction-cost tool. The same structure appears when the “project” is a military commitment rather than a reactor, bridge, or software platform. The public artifact emphasizes near-term necessity and marginal cost. The true burden accumulates across decades and categories. The people paying the integral are not meaningfully present when the commitment is made.
A war-finance claim would look less like institutional dishonesty if the authorizing debate included lifecycle cost ranges, debt-service estimates, veterans’ care obligations, explicit opportunity cost, and scheduled public reauthorization gates tied to updated cumulative cost. The point is not that wartime uncertainty can be eliminated. The point is that uncertainty should expand disclosure before consent, not excuse omission after commitment.
The following caselets are not equivalent to the empirical case studies above. They are included to show transferability of the toolkit across habitats. The enterprise AI caselet is a synthetic/hypothetical pattern. The CPI caselet is an interpretive application to public communication around a real statistical instrument.
This caselet shows the software version of the same protocol. Heavy infrastructure hides uncertainty inside clean baselines; enterprise AI often hides it inside flexible vocabulary. “Agile,” “MVP,” and “pivot” are legitimate engineering concepts, but they become institutional euphemisms when they repeatedly defer the moment at which the engineering ledger must close.
Scenario: A Fortune 500 company announces a multi-million-dollar partnership with an AI vendor to deploy an agentic customer-service system to “revolutionize operations” and reduce headcount burden by 40 percent. This is an interpretive caselet, not a claim about one named company. It shows the software version of the protocol: Agile, MVP, and pivot language are legitimate engineering tools, but they become a Renaming Gambit when they repeatedly delay the point at which the engineering ledger must close.
| Triage Stage | Diagnosis |
|---|---|
| Substrate & Intent | The 40 percent claim is presented like measurement but functions as faith-and-effort: a coordination and signaling device for investors, executives, and internal alignment. |
| Engineering Ledger | A sandboxed demo is treated as evidence of sustained capability. The hidden ledger includes token consumption, data cleaning, retrieval maintenance, cybersecurity, fallback staffing, prompt/version governance, model drift, QA, and exception handling. |
| Temporal Shell Game | When production barriers emerge, the initiative is renamed through Agile/MVP vocabulary: customer-support automation becomes an omnichannel intelligence layer, then a knowledge-copilot platform. The code, vendor, and liabilities continue; the accountability clock resets. |
| Extraction Profile | Support staff, middle managers, customers, and maintenance engineers carry the cost while executives collect short-horizon signaling benefits. |
CPI is a serious statistical instrument. The critique here is not that CPI is fake. The critique is that public communication can treat a narrow technical index as a complete proxy for lived affordability.
| Triage Stage | Diagnosis |
|---|---|
| Substrate & Intent | The statistic is measurement-substrate within its methodology, but rhetoric can stretch it into a broader claim: that household affordability has normalized. |
| Engineering Ledger | Specific methodology can be precise while major lived burdens - housing entry cost, insurance, deductibles, interest rates, quality-adjusted substitutions, and household-specific baskets - remain outside the headline interpretation. |
| Temporal Shell Game | Derivative framing dominates: inflation cooled to a given annual rate. The public, however, lives with the cumulative price level created by prior years. Falling inflation is not falling prices. |
| Extraction Profile | Institutions gain stability from a simple headline; wage earners, retirees, and cash savers live inside the integral. |
The better formulation is not “CPI is fake.” The better formulation is: CPI answers a narrower technical question than the public often thinks it answers, and institutions can benefit when that distinction is blurred.
JWST is included as the framework’s most demanding red-team test, but the red-team lesson is sharper than the earlier version admitted. A red-team that acquits a loved program too easily is not a red-team; it is a character witness. The first JWST appendix properly resisted conviction bias. This version adds the mirror discipline: false-negative bias. The telescope’s scientific glory must not be allowed to purchase innocence for the original sale.
On the surface, the program trips many of the guide’s warning signs: early estimates that grew by roughly an order of magnitude, a launch slip of more than seven years from the confirmed baseline, multiple public rebaselines, a near-cancellation by Congress in 2011, and arguments that the program was too far along to stop. A field guide that convicts every project displaying those surface features is a confirmation-bias machine. But a field guide that exonerates the whole program because the final observatory is magnificent has made the opposite error.
The revised verdict is stage-specific: the sale stage sits in the strategic-misrepresentation zone; the accountability stage is genuinely honest and load-bearing; the technical outcome is a real mission success. JWST is not a clean institutional-dishonesty case as a whole. It is a case where a probably dishonest sale met accountability geometry strong enough to catch, discipline, and recover the program.
The original appendix asked, “Is JWST dishonest?” That was the wrong question. The field guide is temporal. Dishonesty is a property of a lifecycle stage, not automatically of an entire program. A program can be sold with strategic misrepresentation, disciplined through public accountability, and still deliver the promised technical capability. Those facts do not cancel each other. They must be reported separately.
The correct question is: which stage failed, which stage corrected, and which accountability mechanisms remained alive? For JWST, the sale and the correction diverge. The early cost and schedule case look far less innocent than the final mission. The later accountability machinery looks far healthier than the early sale.
The stages are not independent judges casting separate votes. They are a causal chain. In PRA terms, the lifecycle is an event sequence: an initiating event, a set of enabling conditions and propagated consequences, and one or more barriers that may or may not arrest the sequence before it reaches the top event - a materially misleading public artifact governing real money, consent, or program legitimacy.
Grading each stage in isolation hides the thing the framework most wants to see: how the initial distortion propagates, and where it was stopped. For JWST, the useful output is not six separate grades. It is the event sequence.
Sale: strategic misrepresentation -> propagated to laundering and trap -> arrested by live accountability geometry -> residual recovery cost and delay, without permanent public-record distortion.
This reframing also sharpens the Vogtle comparison. Vogtle’s initiating event was also a strategically optimistic sale. The difference is the barrier state. In Vogtle, ratepayer recourse was structurally weak, so the cutset - optimistic sale plus diffuse consequence plus weak recourse plus delayed audit - completed and the top event occurred. In JWST, the equivalent cutset was broken by a high-reliability barrier. Same initiating-event class; different barrier state; opposite top-event outcome.
The payoff is a cleaner output than a grade per stage. Borrowing from PRA importance measures, the Sale carries high risk-achievement worth: it is the event whose distortion most drives the top event across cases. Accountability geometry carries high risk-reduction worth: it is the barrier whose reliability most reduces the probability that a bad sale becomes permanent public distortion.
A program verdict should therefore name the initiating event, trace what it propagated, identify which barrier fired or failed, and state the residual the barrier could not prevent.
Guardrail: this causal chain must not collapse back into convicting the sale for everything. A barrier that genuinely fires produces real recovery, not disguised laundering. The residual cost is the bill for the bad sale; the recovered honesty is a true accountability success. Both are reported. Neither cancels the other.
The gap between early claim and final outcome remains enormous. Early concept costs around the low single billions eventually became roughly $8.8B development through launch, about $9.7B lifecycle through early operations, and roughly $10.8B in 2020 dollars depending on boundary. The confirmed baseline was also breached, leading to a 2011 rebaseline and later 2018 breach/reauthorization. The framework’s job is not to pretend this gap is innocent because the mission succeeded. Its job is to locate the gap in the lifecycle and classify it honestly.
Optimistic Baseline. The early cost estimates did not survive contact with the engineering reality of a first-of-a-kind deployable cryogenic telescope. Hubble’s history, segmented-mirror novelty, sunshield novelty, L2 operations, and no-servicing assumption were knowable reference-class warnings.
Handshake Protocol. Early estimates can be read as the price of admission. A number large enough to reflect the full technical and integration reality might have killed the program before it could prove its scientific case. This is precisely why the sale stage should not be graded as mere innocent uncertainty.
Demonstrate / Operate / Sustain Conflation. Technology maturation milestones were sometimes treated as evidence of integrated-system readiness. Later integration testing exposed problems the maturation milestones had not closed.
Reading-Depth Defense and Specificity Inversion. Early documentation was precise about scientific objectives and less precise about integration risk, reserve adequacy, contractor performance, and total exposure.
Sunk-Cost Pressure. In 2011, the argument against cancellation included the billions already spent. The guide properly flags that move, but the later accountability test determines whether this became colonization or remained a contested political argument.
Renaming / Rebaselining. The program was publicly rebaselined more than once, issuing new cost and schedule envelopes. The key distinction is that the rebaselines remained public and tied to real oversight rather than silently erasing the old baseline.
Public/Private Claim Split. NASA sought additional funding while maintaining public schedule confidence during stressed phases. This marker fires most strongly in the early and middle phases, before the post-Casani correction geometry took hold.
Seven warning signs light up. That does not prove institutional dishonesty. It proves the case must be stage-classified. The sale looks materially worse than the accountability back end. The final telescope does not cleanse the original proposal; the later accountability machinery prevents the original proposal from defining the whole program.
The sale stage is the uncomfortable part. The early low-billion-dollar framing appears to have served as the threshold-crossing number for a program whose real complexity was already visible in the reference class. Hubble had already demonstrated the risks of flagship telescope cost growth. JWST added segmented-mirror deployment, cryogenic operations, a five-layer sunshield, L2 positioning, and no practical servicing assumption. Those were not hidden unknowns.
That does not make the sale fraud. It does not require a secret villain. But under the severity gradient, it is hard to keep the early sale in the ordinary-uncertainty category. The better classification is strategic misrepresentation: a public number shaped to survive an approval gate despite reference-class reasons to doubt it.
Casani’s later finding that the budget had not been adequate and that management, budgeting, and reserves were central problems is the retrospective marker. The public number and the program reality were already diverging at the front end. The diagnostic should say so plainly.
The discriminator is accountability geometry: not whether failure modes appeared, but whether the mechanisms that should catch and discipline failure remained public, load-bearing, and capable of changing the program’s future.
1. Independent review was public and load-bearing. Senator Barbara Mikulski, JWST’s congressional champion, commissioned the Casani Independent Comprehensive Review Panel in 2010. The review was public, identified management and budgeting failure rather than technical impossibility as the central problem, forced a 78% lifecycle cost rebaseline, and triggered structural reorganization. The Farm Filter did not capture the review.
2. Real cancellation off-ramps existed. In July 2011, the House Appropriations Committee voted to cancel JWST. The Senate restored funding after debate. Cancellation was costly, but not fictional; a trapped program does not get cancelled by half the legislature.
3. Congress imposed a binding development cost cap. The $8B development cap was a hard accountability instrument. When the program breached the cap in 2018, NASA had to seek formal reauthorization. The clock did not reset silently.
4. The technical claim held. The mission delivered the core technical and scientific capability it promised: a working L2 infrared observatory with a segmented mirror and sunshield, producing transformative science. This matters because the failure was cost and schedule discipline, not substitution of ceremony for capability.
5. Behavior and public claim eventually converged. After the 2011 rebaseline, NASA’s behavior - separate program office, external review, cost caps, and the 2018 Thomas Young review - increasingly tracked its public commitments rather than contradicting them.
6. Personnel and institutional memory were not laundered. Key actors and institutional seats remained identifiable across the accountability window. The same agency had to answer before the same congressional oversight structure rather than escaping through personnel rotation and silent relabeling.
These markers do not make the original sale clean. They make the later accountability stage clean enough to keep JWST out of the institutional-dishonesty category as a whole. JWST’s contribution to the guide is therefore not exoneration. It is discrimination: dishonest or strategically misrepresented sale conditions can be clawed back by accountability geometry strong enough to bite.
JWST and Vogtle both show optimistic baselines, rebaselines, sunk-cost pressure, and eventual technical delivery. The difference is not that JWST’s front end was clean and Vogtle’s was dirty. The sharper distinction is that JWST’s back end had teeth. The Casani review was public and load-bearing. The cancellation vote was real. The cost cap bound. The breach required reauthorization. The technical claim held without quietly redefining the mission downward.
Vogtle, by contrast, combined a weak public-consent structure with rate-recovery geometry that diffused consequences onto ratepayers who had limited practical off-ramp. That is why the same surface signs lead to different stage verdicts. The guide should not ask whether a whole program is beloved, useful, clean-energy, beautiful, or scientifically glorious. It should ask where the sale diverged from known reality and whether later institutions forced the truth back into the public record.
A. Add false-negative discipline. The guide already guards against convicting every failure. It must also guard against acquitting a successful or admired program because the outcome was valuable. Final success does not prove the sale was honest.
B. Output verdicts by stage. The diagnostic should classify Setup, Sale, Laundering, Trap, Accountability, and Countermeasure separately. “Is the program dishonest?” is usually too blunt. “Which lifecycle stage was dishonest or strategically misrepresented, and which stage corrected it?” is the better question.
C. Preserve the Severity Gradient. JWST’s sale sits closer to strategic misrepresentation than ordinary uncertainty. Its accountability stage sits closer to honest correction than institutional dishonesty. The program-level result is therefore mixed, not binary.
D. Keep the negative control, but make it harder. JWST remains a red-team case because it prevents the framework from convicting on surface specimens alone. The revised red-team is stronger because it also prevents the analyst from letting admiration for the final product erase front-end distortion.
E. Treat honesty markers as correction markers, not absolution markers. A public review, binding cap, or real cancellation vote can discipline a bad sale. It does not retroactively make the sale good.
F. Model stage verdicts as propagating sequences, not independent cells. The diagnostic should identify the initiating event, the propagated consequences, the barrier state, and the residual consequence. This prevents double-counting common-cause specimens while preserving the central insight: a bad sale can be real, a later correction can be real, and neither fact erases the other.
The earlier case studies ask what would have made the case look more like ordinary uncertainty. Here the inversion must be staged.
The sale would look more innocent if the early public estimate had been published with explicit reference-class warnings, wide cost ranges, non-serviceability risk, reserve uncertainty, and a plain statement that the number was conceptual rather than a reliable program baseline.
The accountability stage would look dishonest if the Casani report had been suppressed or filtered before reaching Congress; if the 2011 cancellation vote had been a managed performance with no possibility of execution; if the 2018 cost-cap breach had been concealed or quietly rebaselined without reauthorization; if the science claims had been quietly downgraded while marketing language continued to promise transformative capability; or if NASA’s internal schedules, contingencies, and contractor management revealed persistent disbelief in the public timeline while that public timeline continued to be defended.
Those later patterns did not hold. That is the test. But their absence does not erase the front-end sale problem. It shows that accountability can recover a program that was not honestly sold.
| Dimension | Public / Early Case | Later Reality |
|---|---|---|
| Cost - early concept | ~$1B decadal survey estimate in the early 2000s; roughly $1-3.5B as the program matured through 2005. | ~$8.8B development through launch; roughly $9.7B lifecycle through five years of operations; about $10.8B in 2020 dollars. |
| Cost - confirmed baseline | About $5B at the 2008/2009 confirmed baseline. | A 78% lifecycle cost increase following the 2011 rebaseline. |
| Schedule | Launch in the mid-2010s at the confirmed baseline. | Launched December 25, 2021 - more than seven years late from the confirmed baseline. |
| Technical claim | First-of-a-kind deployable cryogenic infrared telescope at L2 with an unprecedented sunshield and segmented mirror. | Delivered as advertised; first science images released July 2022; transformative scientific output since. |
| Accountability | Independent reviews, congressional cost caps, public rebaselines. | Casani review in 2010, Thomas Young review in 2018, congressional development cap, and formal public reauthorization after breach. |
| Section | Entry |
|---|---|
| Claim at Approval | Operating posture rather than a single approval number: New Glenn was positioned as flight-ready and on a near-term cadence to begin deploying Amazon's Leo broadband constellation, after returning to operations from an earlier 2026 anomaly. |
| Known or Knowable Baseline | First-of-a-kind heavy-lift reference-class risk; an unresolved-or-recent engine-related anomaly from April 2026; documented schedule pressure (constellation deployment deadlines, Artemis lander timeline, competitive framing); a single launch pad with no operational backup. |
| Lifecycle Map | Sale: possible Optimistic Baseline (cadence formed under schedule pressure). Laundering: indeterminate pending root cause. Trap: not yet present - watch the rebuild justification. Accountability escape: the live question - external investigator may not engage. |
| Consequence Map | Gain from the posture: schedule credibility against FCC milestones and a competitor. Cost of the miss: vehicle and pad lost, constellation timeline slipped, Artemis lunar-lander schedule exposed. Recourse: a customer/regulator/contractor question still open. |
| Falsification Test | The dishonesty reading collapses if the April anomaly was fully root-caused and closed before the hot fire, an independent investigation with publication rights engages, and no suppressed-dissent pattern exists in the current program. |
| Lessons for the Guide | Motivates an explicit Observability Limit discipline: a live case where the verdict is necessarily an interval because the evidence is endogenous to the subject and the root cause is unknown. |
---
On the night of May 28, 2026, a Blue Origin New Glenn rocket was destroyed during a static-fire (hot-fire) test at Launch Complex 36, Cape Canaveral, igniting in a fireball that also caused severe damage to the launch infrastructure. No one was injured, and no customer payload was aboard. As of this writing the root cause has not been established; the company has said publicly that it is too early to know.
This case is included not to assign a verdict but to show what disciplined analysis looks like when a verdict is not yet available. It is a live event, the evidence is largely held by the institution under examination, and several lifecycle stages cannot be resolved without facts not yet in the public record. Under the Burden-of-Proof Rule, that combination forbids a finding of institutional dishonesty. It equally forbids a clean acquittal. The honest output is an interval whose width is set by how little can currently be observed - and the case is here to demonstrate that move.
The case is not anti-Blue-Origin and not anti-spaceflight. Developing heavy-lift launch capability is genuinely hard, and catastrophic ground-test failures occur in honest programs. The case is anti-premature-verdict.
| Dimension | Public / Operating Posture | Outcome / Open Question |
|---|---|---|
| Launch cadence | Positioned to begin deploying roughly fifty Amazon Leo satellites per flight on a near-term schedule, against FCC deployment milestones and a direct competitor. | Vehicle and pad destroyed in a ground test; cadence disrupted indefinitely; constellation timeline pressure intensified. |
| Return to readiness | The rocket had returned to operations after an April 2026 anomaly (engine-related; a satellite was left in the wrong orbit) and the subsequent grounding was lifted. | A fueled hot-fire ended in vehicle loss weeks after return. Whether the April root cause was fully closed before this test is the central open question. |
| Engine maturity | Seven BE-4 first-stage engines treated as ready for a hot-fire test. | One unconfirmed early account describes a pressure excursion in engine components beyond design tolerance. Root cause is not established and should be treated as unverified. |
| Infrastructure resilience | LC-36 operated as the launch site for New Glenn. | The sole New Glenn pad was heavily damaged; a second Cape pad and a Vandenberg site are reportedly not yet underway; repairs are expected to take months, with no operational backup. |
| Lifecycle Stage | New Glenn Pattern |
|---|---|
| The Sale | Possible Optimistic Baseline. The cadence and readiness posture formed under heavy schedule pressure - constellation deployment deadlines, the Artemis lunar-lander timeline, and an openly competitive framing. Whether that optimism crossed into misrepresentation is undetermined and may stay so. |
| The Laundering | Indeterminate. A laundering finding would require evidence that a known, specific concern was filtered out of the readiness decision. A historical allegation describing exactly such a mechanism exists (see E.4 and E.7) but dates to prior leadership and is not established for this program. |
| The Trap | Not yet present. There is no visible Sunk-Cost Colonization at this stage. The variable to watch is the rebuild justification. The qualitative flag is a shift from physical booster qualification to protection of Amazon Leo/FCC deployment deadlines or other downstream commitments. Once the argument becomes “we must rebuild quickly because the constellation deadline, customer cadence, or Artemis dependency cannot slip,” the program has crossed into the Trap stage: past commitments begin governing the go/no-go logic more than the physical readiness of the booster and pad. |
| Accountability | The live question, and the highest-information stage available. Reporting indicates the static fire fell outside the scope of FAA-licensed activity, which raises the possibility that the independent investigator with publication rights does not formally engage and the inquiry proceeds largely inside the company. Counter-weight: early public statements from the company and from NASA leadership named the program impact as an open question rather than pre-reassuring it away - a marker on the honest side of the ledger. |
The diagnostic core of this case is not the explosion. It is the absence - so far - of any contemporary engineer publicly warning that this vehicle was unsafe to test. A naive reading treats that silence as reassuring. The guide treats it as nearly uninformative, and explaining why is the case's contribution.
A bad outcome occurred. Consider the space of what the silence could mean:
The silence cannot distinguish these. Its likelihood ratio between the best reading (unforeseeable) and the worst (warned-and-overridden) is near one; it discriminates almost nothing. This is the Observability Limit in operation: the analyst is interrogating the single variable the institution most controls - whether anyone inside was worried - and that variable returns the same null whether the program was honest-and-unlucky or rotten-and-deferential.
The disciplined response is to stop interrogating the silence and interrogate the architecture that would create or suppress a warning, because that architecture is more observable than anyone's private misgivings. Three observables carry the load:
1. Precursor closure: was the April engine anomaly fully root-caused and closed before return to fueled hot-fire? 2. Dissent channel: could an engineer halt the test without career penalty? 3. Barrier teeth: if static fire sits outside licensed activity, does any external body with publication rights investigate, or does the company investigate itself?
A critic should be invited to falsify any dishonesty reading. This case moves toward ordinary engineering risk, honestly borne if: the April anomaly was demonstrably root-caused and closed before the hot fire; an independent investigation with genuine publication rights engages and its findings are released; the root cause is disclosed rather than absorbed; and no pattern of retaliated-against dissent exists in the current program.
It moves toward institutional dishonesty only if specific facts appear: a documented return to fueled test before precursor closure under recorded schedule pressure; an inquiry kept entirely internal with no external publication; and a recurrence of the suppressed-dissent pattern alleged historically.
The essential discipline: at the time of writing, the test cannot be run, because the facts it requires are not yet observable. Recording that inability honestly - rather than resolving it with a confident verdict in either direction - is the correct output. The verdict is an interval, and the interval is wide because the subject controls the evidence.
This case pressure-tests the guide on an axis JWST does not. JWST disciplines the false positive: do not convict a program merely because surface specimens light up. New Glenn disciplines the observability limit: do not convict or acquit when the evidence is endogenous to the subject and the root cause is unknown - and say so explicitly, with the verdict expressed as an interval whose width encodes the blindness.
Three specific contributions:
A further design implication follows: the licensed-activity gap is also a response-design problem. If a safety or accountability barrier exists on paper but does not fire when the failure occurs outside its formal scope, the barrier is not a barrier for that event. The field guide should treat barrier scope as part of the accountability architecture, not as an administrative footnote.
Multiple outlets (Spaceflight Now, CBS News, Scientific American, CNBC) document the May 28, 2026 static-fire failure at Launch Complex 36, the destruction of the vehicle and damage to the pad, the absence of injuries and of an onboard payload, and the company's statement that the root cause was not yet known.
Reporting on the company's competitive and schedule context (Amazon Leo deployment against FCC milestones, the Artemis lunar-lander dependency, and the single-pad infrastructure posture with a second Cape pad and a Vandenberg site not yet underway) is documented in trade and financial coverage of the event.
The early account of a pressure excursion in engine components exceeding design tolerance appears in secondary reporting and should be treated as an unverified reconstruction, not an established root cause.
The allegation that schedule pressure drove safety-process violations is drawn from a 2023 wrongful-termination lawsuit filed by a former BE-4 program manager. It is an unproven allegation, it concerns the engine family rather than this specific event, and it dates to a prior chief executive's tenure. It is included as a structural prior to be checked, not as a finding.
The regulator's characterization of the static fire as outside the scope of its licensed activity is drawn from the regulator's statement to CNBC and is the basis for the open question about whether an external investigation with publication rights engages.
Claim at approval: Public-health authorities attributed the microcephaly surge primarily to Zika virus and publicly treated the pyriproxyfen hypothesis as unsupported by available evidence.
Known or knowable baseline: Zika had high-specificity affirmative evidence, including viral material in affected fetal/infant tissue and later animal and epidemiological support. Pyriproxyfen had a lower-specificity but non-random trigger: a developmental-disruption mode of action in insects, a plausible retinoid/thyroid pathway concern, secondary reports of mammalian developmental findings in rat pups, and unprecedented drinking-water use with uncertain delivered exposure.
Lifecycle map: Two maps must be kept separate. Causal map: agent identification under emergency uncertainty. Institutional map: public reassurance, evidentiary disclosure, follow-up burden, and whether the rival hypothesis was suppressed or allowed to test itself.
Consequence map: The public-health gain from rapid closure was mosquito-control continuity during an epidemic. The cost of being wrong would have been continued exposure to a possible developmental toxicant. The accountability geometry partly fired: the hypothesis was aired, use was suspended in at least one jurisdiction, and independent testing and case-control work followed.
Falsification test: The over-reassurance concern weakens if public communications show calibrated uncertainty, the rat-pup signal was disclosed and directly addressed, and adequately powered mammalian developmental follow-up was commissioned. The pesticide-causation claim weakens further on negative human-relevant developmental tests and case-control evidence showing no association.
In 2015-2016, northeastern Brazil recorded a surge in microcephaly and related neurodevelopmental defects. Public-health authorities attributed the surge to a concurrent Zika epidemic. A rival hypothesis proposed that pyriproxyfen, an insect-growth-regulator larvicide used in some drinking-water containers, was the cause or a cofactor.
The appendix asks two questions, not one: did pyriproxyfen cause or contribute to the harm; and did institutions handle the uncertainty honestly? These questions can have different answers. If the pesticide did not cause the surge, it does not automatically follow that the institutions communicated perfectly. If institutions over-reassured, it does not follow that the pesticide caused the surge. The guide must keep the questions apart.
This appendix is not a finding that pyriproxyfen caused the Brazilian microcephaly surge. The evidentiary balance supports Zika as the primary cause. The appendix asks a separate institutional question: whether public-health authorities communicated residual pesticide uncertainty with appropriate calibration.
The public record narrows the indictment. Brazil, WHO/PAHO, and CDC generally used no-evidence, no-link, or no-scientific-basis language rather than claiming logical impossibility. That matters. The record does not support a strong institutional-dishonesty finding.
Brazil Ministry of Health, February 2016: public statements rejected the pyriproxyfen association on the ground that no epidemiological study proved it and that areas without pyriproxyfen also reported microcephaly. Field-guide read: dismissive, but framed as absence of supporting evidence rather than proof of impossibility.
WHO/PAHO, March 2016 and later: public messaging described no evidence that pyriproxyfen causes microcephaly and relied on toxicological review, regulatory approval, and exposure arguments. Field-guide read: mostly calibrated as no evidence, but practical reassurance language risks compressing residual developmental uncertainty.
CDC key messages, 2016: CDC stated that pyriproxyfen had not been linked with microcephaly, that WHO had approved it for mosquito control, and that pyriproxyfen exposure would not explain Zika virus detected in infant brains. Field-guide read: comparatively calibrated, and anchored in affirmative Zika evidence.
Later evidence: Pernambuco case-control work supported the Zika association and did not find pyriproxyfen use associated with microcephaly. Field-guide read: later evidence cuts against pyriproxyfen as a primary cause and shrinks any institutional-conduct claim toward over-reassurance, not dishonesty.
The institutional issue is not that authorities failed to prove an infinite negative. No public-health body can prove every suspected agent harmless in all possible exposure windows. The issue is narrower and more serious: pyriproxyfen was not a random suspect.
Pyriproxyfen is a juvenile-hormone analog / insect growth regulator. Its intended function is to disrupt developmental maturation in insects. Critics identified plausible developmental-pathway concerns, including retinoid/thyroid-system arguments. Secondary reviews of the manufacturer's toxicology data reported mammalian developmental findings including low brain weight and arhinencephaly in exposed rat pups. Those claims remain contested and should be verified against primary toxicology reports before any safety finding is treated as settled.
That combination did not prove causation in humans. But it did create a reasonable trigger for follow-up work. A parent deciding whether to accept pyriproxyfen-treated water around children would reasonably ask for more than no epidemiological proof of association; they would ask whether the specific mammalian developmental signal had been followed up with targeted, adequately powered developmental testing under plausible exposure windows.
Public reassurance was cheaper than actually closing the toxicology question. That does not make the reassurance fraudulent. It does mean the confidence borrowed from registration status and no-evidence phrasing should not be mistaken for experimental closure of the specific developmental concern.
The opposing case is stronger on the causal question. Zika viral material was identified in damaged fetal or infant tissue; later animal models and clinical characterization supported congenital Zika syndrome; and the timing and geography of the epidemic fit Zika better than they fit a larvicide-use map.
The pesticide hypothesis also has specificity problems. Pyriproxyfen was used in many places without comparable microcephaly surges. The reported rat finding was not a human-relevant demonstration and, as commonly summarized, involved high experimental doses and non-recurrence at a higher dose. A zebrafish study and later epidemiological work cut against the pesticide-causation claim. The result is asymmetry, not balance.
Causal verdict: Zika is strongly supported as the primary cause of the Brazilian microcephaly surge. Pyriproxyfen as primary cause is unsupported. Pyriproxyfen as a cofactor remains speculative and unproven.
Institutional-conduct verdict: the authorities were probably justified in rejecting the pesticide-cover-up theory, but not justified in communicating practical reassurance without visibly discharging the triggered toxicology burden. The failure is not proven causation; it is unearned certainty.
This is not institutional dishonesty and not fraud. It is best classified as possible over-reassurance / unearned practical closure under emergency conditions, observability-limited and source-dependent. The safety dossier was partly controlled by interested parties, and the public record does not show the specific mammalian-signal concern being closed by targeted follow-up before reassurance was communicated.
The reader should not take this appendix to mean pyriproxyfen caused the outbreak. The reader also should not take it to mean the compound was obviously safe around children. The correct output is narrower: the evidence did not prove pyriproxyfen caused the harm, but the evidence also did not justify the level of reassurance the public received.
The intuitive grievance - they never proved it was not involved - is a trap if taken literally. Proving a negative to completion is rarely possible, and the demand can be levied against any institution, about any agent, indefinitely. Used that way, it becomes a worldview-installation engine.
The defensible reformulation is narrower and stronger: identify a specific signal that triggered a duty to investigate further, then ask whether that burden was discharged before closure was communicated. Here, the alleged mammalian developmental signal plus a plausible developmental mechanism created the trigger. The issue is not failure to prove infinite safety; it is whether certainty outran the follow-up that the trigger reasonably demanded.
Public-claim record: CDC Zika key messages; CDC 2016 causal statement; WHO/PAHO Zika and pyriproxyfen communications; Brazil Ministry of Health statements reported in contemporary coverage. Scientific record: Pernambuco case-control work in The Lancet Infectious Diseases; ecological analyses of pyriproxyfen use and microcephaly prevalence; zebrafish disconfirmation literature; NECSI/Bar-Yam and related papers presenting the pyriproxyfen hypothesis and toxicology concerns. Toxicological specifics should be verified against primary study reports before any final safety conclusion is drawn.
Institutional dishonesty is not simply the presence of bad people inside large systems. It is what happens when authority, consequence, language, and publication rights are arranged so that honest analysis cannot survive contact with approval.
The institution may contain honest people. It may produce honest internal work. It may even believe, at each layer, that it is behaving responsibly. Yet the public artifact can still be dishonest.
Where the tests fail, the polite vocabulary should end. Where the tests pass, the analyst should stand down. The point is not to call every failure dishonest. The point is to preserve the difference between failure honestly disclosed and failure made inevitable by a public artifact that never deserved belief.