·

Who Decides When Frontier AI Is Ready?

The most consequential decision in artificial intelligence governance currently happens inside the institution that benefits from saying yes.

A developer decides that a frontier model is ready. The company has already chosen which capabilities to test, what counts as sufficient evidence, which uncertainties are acceptable, and who has access to the results. By the time the system reaches the public, the investment has been made, the infrastructure has been built, and competitors are preparing their response. Public oversight enters after the decision has hardened.

This does not mean frontier developers are indifferent to safety. They face real security threats, investor expectations, global competition, and the possibility that delay will transfer influence to actors with fewer safeguards. Governments see scientific capacity, economic growth, national defense, and strategic infrastructure. Communities see jobs and revenue beside rising demands on energy, water, land, and public services.

Each actor can behave rationally within a role while the whole system becomes less capable of seeing itself.

The Argument Begins With the Wrong Choice

Public debate often frames artificial intelligence as a choice between stopping a dangerous technology and allowing innovation to proceed. One side protects safety, democratic authority, work, privacy, and the right to remain free from risks imposed by private institutions. The other protects discovery, medicine, accessibility, economic movement, national security, and forms of problem-solving that may exceed unaided human capacity.

Both sides protect something real. Neither becomes adequate by defeating the other.

A general prohibition would concentrate political authority, delay beneficial uses, protect established institutions from challenge, and transfer development to jurisdictions willing to accept greater risk. Unrestricted acceleration would allow the institutions creating frontier capabilities to define readiness, acceptable evidence, and public exposure within the same decision.

The conflict becomes more precise when we separate current systems from imagined futures. Existing large language models are artificial intelligence. They can use tools, write code, sustain longer sequences of work, and assist parts of research and development. They are not established examples of artificial general intelligence or superintelligence. Fully autonomous recursive self-improvement is not an established present condition, although AI systems are increasingly contributing to the work used to develop later systems.

We do not need to settle the future of machine intelligence before governing the authority already being created. A system does not need general intelligence or independent intention to reorganize labor, infrastructure, information, security, or access to public goods. It needs capability, scale, and permission.

The immediate problem is smaller than the ultimate fear and more available to intervention. Society lacks a dependable decision between a developer’s internal conclusion that a frontier system is ready and the public deployment that distributes its consequences.

Capability Is Easier to Measure Than Consequence

Technical capability produces legible evidence. A model can complete a benchmark, reproduce an experiment, discover a vulnerability, or carry a task for a measurable length of time. The consequences of that capacity move through systems that use different forms of evidence and answer to different authorities.

An energy regulator sees grid demand. A labor agency sees displacement and transition. A security evaluator sees cyber or biological capability. A local government sees land use, tax revenue, and water demand. A developer sees performance, market position, and product risk. Each institution holds part of the decision. No institution necessarily holds the relationship among them.

Fragmentation creates a predictable result. Model development and material infrastructure move through separate approval processes. Security evidence remains confidential. Ecological effects arrive through facilities rather than software. Labor consequences emerge after adoption. Public agencies respond within their jurisdiction, even when the risk exists in the connections among jurisdictions.

Confidentiality carries its own dialectic. Frontier developers possess information that should not be released publicly: model weights, vulnerabilities, dangerous evaluation details, personal information, and legitimate trade secrets. Total disclosure can create the harm oversight is meant to prevent. Total secrecy asks the public to accept conclusions it cannot evaluate.

Once confidentiality contains both dangerous information and inconvenient evidence, outsiders cannot tell which one is being protected. Mistrust produces demands for broader disclosure. Broader demands make developers more protective. The system becomes less capable of the differentiated transparency it needs.

The public is left carrying uncertainty without authority. The people deciding to proceed retain the evidence without carrying the full consequence.

A Decision Before Deployment

California has already built part of the necessary structure. Its frontier AI law requires large developers to publish safety frameworks, report critical safety incidents, protect whistleblowers, and remain accountable to enforcement by the Attorney General. These requirements make declared practices and serious failures more visible. They do not yet create an integrated predeployment judgment about whether a consequential system should be released, limited, or held.

A Frontier AI Systemic Impact Case could fill that gap.

The requirement would apply at a capability threshold, not to every model update or small developer. A covered company preparing a major public release or a material expansion of automated AI research would have to present its case before deployment. The burden would rise with the capacity to create distributed consequences.

The case would contain two records. A confidential technical annex would give qualified independent evaluators access to methods, results, security controls, internal disagreement, material demands, and proposed stopping conditions. A public impact statement would identify what was tested, what remains unknown, which systems and communities carry the burden, who authorized the decision, how the deployment will be monitored, and what would cause it to narrow or stop.

This structure preserves a distinction that public debate often loses. Accountability does not require publishing information that creates a specific security risk. Security does not justify concealing uncertainty, dissent, burden, or the absence of a stopping condition.

Independent reviewers would test the developer’s claims in a secure environment. Technical evaluation would sit beside review of energy, water, infrastructure, labor, and community effects. No single benchmark would become proof of safety. No generalized concern would become an indefinite veto.

The final decision would allow more than approval or prohibition. A system could be approved under its proposed conditions, approved within a narrower scope, released as a limited pilot, or deferred until a named evidentiary condition is met. Continued deployment would remain connected to incident reporting, significant capability changes, material-demand variance, and the performance of safeguards after release.

This is the difference between regulation as a barrier and regulation as a feedback loop. A barrier decides once. A feedback loop preserves the capacity to observe and revise.

The Cost Is Part of the Decision

A predeployment gate would impose real losses. Large developers would surrender some speed, secrecy, surprise advantage, and unilateral control. Investors would lose certainty over release timing. Engineers would explain systems to public institutions that may understand less of the technology than they do. Government officials would inherit responsibility for approving deployments whose consequences cannot be fully known.

Delay could transfer development to another jurisdiction. Confidential review could expose sensitive information. Filing costs could strengthen incumbent firms that can maintain permanent legal and compliance teams. An ecological assessment could become another requirement that large companies absorb and smaller competitors cannot.

These losses do not invalidate the mechanism. They identify who currently benefits from the absence of it.

The current system also imposes costs. Workers adapt after deployment. Utilities build around demand they did not shape. Communities negotiate infrastructure after commercial expectations have formed. The public receives safety claims without access to the evidence beneath them. When harm appears, responsibility is distributed across institutions that can each explain why the decisive choice belonged somewhere else.

The proposed change redistributes uncertainty. The public carries less of it without authority. Developers, investors, evaluators, and officials carry more of it with responsibility.

That is a genuine sacrifice, not an administrative inconvenience.

Oversight Can Become Its Own Form of Theater

The mechanism contains the pattern it is meant to correct. A process designed to distribute accountability can consolidate power among the institutions capable of satisfying it. A public impact statement can become safety marketing. Accredited evaluators can become dependent on developer access and fees. An approval certificate can teach the public to outsource judgment to a process it cannot see.

The existence of oversight can then become evidence of safety, even when the oversight has stopped producing information that changes decisions.

This is why the mechanism should not begin at full scale. A bounded procurement pilot could test whether the review finishes on time, protects confidential information, identifies evidence gaps the developer’s existing framework missed, and produces conditions useful to a real deployment decision. The gate should expand only if the surrounding institutions demonstrate the capacity to operate it.

The review also needs its own stopping condition. If it repeatedly produces delay without decision-relevant findings, strengthens market concentration, or exposes protected information, the process must narrow, pause, or be redesigned. Accountability cannot mean that the technology remains revisable while the regulator becomes permanent.

The solution remains partial. California cannot resolve international competition, federal authority, labor distribution, data-center infrastructure, or the future direction of artificial intelligence through one state process. It can make one consequential decision more visible before that decision becomes difficult to reverse.

Readiness Is a Relationship

The usual question is whether a frontier model is ready. That places readiness inside the model and the company that built it.

A more complete question includes the systems that must receive it. Are evaluators capable of testing its claims? Can infrastructure absorb its material demand? Can affected institutions observe what changes? Can a deployment be narrowed when evidence shifts? Is authority clear enough that someone remains responsible after release?

Innovation and accountability are often treated as competing speeds. The deeper relationship is capacity and response. As the capacity of a system grows, the surrounding capacity to observe, contain, and revise it must grow with it. When one accelerates while the other remains fragmented, private readiness becomes public exposure.

We do not need certainty about superintelligence to recognize that structure. We need enough differentiation to separate present capability from imagined destiny, enough deconstruction to see why rational actors keep accelerating, and enough dialectical maturity to preserve innovation without assigning its uncertainty to people outside the decision.

A frontier system is ready only when the systems carrying its consequences are ready to respond.

═══════════════════════════════════════════════════════════════
DIALECTIC AND DECONSTRUCTION SOLUTIONS (DDS)
BLUEPRINT
═══════════════════════════════════════════════════════════════

Problem: Frontier artificial intelligence is advancing without an enforceable predeployment process that connects capability growth to the ecological, social, and institutional systems that must absorb its consequences.

───────────────────────────────────────────────────────────────

PHASE 1: PROBLEM FRAMING

The Umbrella Problem

Human institutions do not yet have a reliable way to keep increasingly capable artificial intelligence accountable to the material, ecological, relational, and civic systems that make its development possible.

Factual Distinctions

Current large language models are artificial intelligence systems. They generate outputs from learned statistical structure and can increasingly use tools, sustain multistep work, and participate in research and software development. They are not thereby established examples of artificial general intelligence, a term without a single agreed technical threshold that usually refers to broad competence across domains comparable to or exceeding human capability.

Superintelligence remains a hypothetical condition in which machine capability substantially exceeds the strongest human capability across most consequential domains. Automated AI research is narrower: current systems can already perform parts of the work used to evaluate and develop later systems. Recursive self-improvement would require an AI system to drive successive generations of increasingly capable AI with diminishing human direction. OpenAI states that fully autonomous recursive self-improvement is not happening today. Anthropic likewise states that it has not arrived and is not inevitable, while reporting evidence that AI is accelerating parts of AI development. These are reasons to measure the pathway, not evidence of an inevitable exponential movement toward infinity. OpenAI Anthropic

The ontological question remains outside the reach of these definitions. We do not know whether direction originates in intelligence, embodiment, relationship, material constraint, consciousness, or some structure in which those distinctions participate. This blueprint treats that uncertainty as a design condition. It does not assume that greater intelligence naturally becomes ethical, destructive, self-preserving, ecologically aware, or aligned with human intention.

The Multiple Drivers

  • Missing predeployment visibility: No shared process makes frontier capability, automated AI research, ecological demand, downstream burden, and stopping conditions legible before deployment.
  • Competitive acceleration: Companies and states experience restraint as a possible transfer of advantage to less cautious rivals.
  • Fragmented authority: Model development, energy infrastructure, national security, labor, consumer protection, and environmental effects are governed by different institutions with incomplete jurisdiction.
  • Externalized material cost: Energy, water, grid infrastructure, mineral demand, emissions, and local land-use burdens can be separated from the product decisions creating them.
  • Unresolved direction and alignment: Human institutions are attempting to direct systems whose future capacities are uncertain while our own values, incentives, and definitions of public benefit remain contested.

This Blueprint Addresses

This blueprint addresses missing predeployment visibility by creating a California Frontier AI Systemic Impact Case and Conditional Deployment Gate. The mechanism would extend California’s existing frontier-AI framework so that covered developers must connect technical capability to material demand, affected communities, institutional readiness, and enforceable stopping conditions before a major deployment or a material increase in autonomous AI research.

Remaining Components

  • International coordination and jurisdictional competition
  • Market incentives, labor transition, ownership, and distribution of AI-created value
  • Energy, water, mineral, and grid policy for data-center expansion
  • The unresolved scientific and philosophical problem of alignment, agency, consciousness, and ontological direction

Bounded Ambition Note: This blueprint addresses missing predeployment visibility. It does not attempt to resolve international competition, infrastructure externalities, labor distribution, or the origin and direction of intelligence, which require separate interventions.

───────────────────────────────────────────────────────────────

PHASE 2: DECONSTRUCTION

The Surface Symptom

Public discussion alternates between promises that AI will solve problems beyond human capacity and warnings that it could destabilize economic, political, ecological, or military systems. Companies publish safety frameworks, governments develop standards, and evaluators test selected capabilities. The public still lacks a dependable way to know what was tested, which consequences were considered, who accepted the remaining risk, and what would cause a deployment to stop.

The False Start

The problem is commonly framed as a choice between stopping dangerous technology and allowing innovation to proceed. That frame treats acceleration and prohibition as the only coherent positions and turns uncertainty into a contest of confidence.

The Compassionate Reality

The institutions moving quickly are responding to real pressures. Developers face competitors, investor expectations, security concerns, rapidly changing capabilities, and the possibility that delay transfers influence to actors with fewer safeguards. Governments see economic growth, scientific capacity, national security, and control of strategic infrastructure. Communities see jobs and tax revenue alongside rising energy demand, water use, land conversion, and the possibility of costs they did not choose. The difficulty is not a lack of intelligence or concern. It is that each actor can behave rationally within a role while the whole system becomes less capable of seeing itself.

The Upstream Drivers

  • Race-Condition Development
    • Actor(s): Frontier AI developers, investors, cloud providers, and national governments
    • Incentive/Constraint: First-mover advantage, capital expectations, talent competition, national-security pressure, and fear that unilateral restraint benefits rivals
    • Behavior: Compress evaluation timelines, keep evidence private, deploy capability before external institutions can assess downstream effects, and treat delay as lost position
    • Loop: Faster releases reset competitive expectations. Competitors accelerate in response, shortening the next evaluation window and making restraint increasingly expensive.
  • Fragmented Measurement and Authority
    • Actor(s): Standards bodies, safety institutes, sector regulators, energy agencies, state governments, and international institutions
    • Incentive/Constraint: Each institution has limited jurisdiction, different legal authority, and responsibility for only part of the risk landscape
    • Behavior: Measure cybersecurity, biological risk, energy demand, labor effects, privacy, or consumer harm separately, without a shared case connecting them
    • Loop: Fragmented evidence prevents a whole-system decision. The absence of an integrated decision reinforces the belief that integration is administratively unrealistic.
  • Confidentiality as Both Protection and Blind Spot
    • Actor(s): Developers, security teams, regulators, independent evaluators, and the public
    • Incentive/Constraint: Model weights, evaluation results, security vulnerabilities, trade secrets, and research methods can create real danger or competitive loss when disclosed
    • Behavior: Developers restrict access; regulators rely on voluntary summaries; the public receives conclusions without the evidence required to evaluate them
    • Loop: Limited visibility increases mistrust and political pressure for blunt disclosure. That pressure makes firms more protective, further reducing the trust needed for differentiated oversight.
  • Material Costs Outside the Model Decision
    • Actor(s): AI developers, data-center operators, utilities, grid planners, local governments, ratepayers, and communities near large facilities
    • Incentive/Constraint: Developers need compute; utilities seek load growth and infrastructure recovery; local governments seek investment and revenue; affected residents have limited access to model-development decisions
    • Behavior: Model scaling and facility approval proceed through separate processes, allowing product decisions to be detached from electricity, water, transmission, and land-use consequences
    • Loop: Infrastructure expansion lowers the immediate constraint on further scaling. New capacity creates pressure to keep facilities utilized, which reinforces demand for more compute-intensive development.
  • Uncertainty Converted Into Ideology
    • Actor(s): Developers, policymakers, researchers, journalists, advocates, investors, and the public
    • Incentive/Constraint: Uncertainty is difficult to communicate and politically costly to hold. Institutions gain attention and legitimacy by offering confidence.
    • Behavior: Speculative futures are presented as inevitabilities or dismissed as fantasy; present harms are conflated with hypothetical extinction; beneficial possibilities are treated as proof of safety
    • Loop: Polarized claims reduce shared measurement. Weak shared measurement increases uncertainty, which creates more room for polarized claims.

The Entry Point

The entry point is the gap between a developer’s internal decision that a frontier system is ready and society’s later discovery of what that deployment changed. California already requires large frontier developers to publish safety frameworks, report critical safety incidents, protect whistleblowers, and answer to enforcement by the Attorney General under SB 53. That creates a beam capable of carrying more weight, but the current structure remains oriented toward declared frameworks and post-incident accountability. A predeployment Systemic Impact Case would connect capability, material demand, burden distribution, uncertainty, and stopping conditions before those consequences become distributed across institutions that cannot reconstruct the original decision.

───────────────────────────────────────────────────────────────

PHASE 3: DIALECTICS

The Core Tensions

Primary: Intelligence Capacity ↔ Ecological Embeddedness

Secondary: Innovation ↔ Accountable Continuity

Secondary: Transparency ↔ Security and Privacy

Secondary: Urgency ↔ Sustainability

The Weighting

  • Intelligence Capacity ↔ Ecological Embeddedness
    • Current State: 85% capacity expansion / 15% embedded accountability
    • Target State: 65% capacity expansion / 35% embedded accountability
  • Innovation ↔ Accountable Continuity
    • Current State: 80% innovation / 20% continuity
    • Target State: 65% innovation / 35% continuity
  • Transparency ↔ Security and Privacy
    • Current State: 25% public transparency / 75% confidentiality
    • Target State: 45% public transparency / 55% confidentiality
  • Urgency ↔ Sustainability
    • Current State: 75% immediate development / 25% durable absorption
    • Target State: 55% immediate development / 45% durable absorption

Who Benefits: The public, downstream developers, workers, communities carrying infrastructure burdens, regulators, and frontier developers whose long-term legitimacy depends on credible restraint.

Who Bears Cost: Large frontier developers, cloud providers, investors, fast-following deployers, and agencies required to build technical competence.

What Is Sacrificed: Some product secrecy, unilateral control, surprise advantage, and release speed; a covered deployment may be delayed by 30 to 90 days or restricted to a smaller environment.

Redistributed Emotional Burden: The public currently carries uncertainty without authority. The rebalancing moves more uncertainty onto executives, investors, and public officials, who must tolerate delay, disclose limits, and accept responsibility for a decision that cannot be made risk-free.

Intelligence Capacity ↔ Ecological Embeddedness

Capacity protects discovery, adaptation, defense, medicine, coordination, and the possibility of solving problems that exceed unaided human cognition. Embeddedness protects the conditions under which any intelligence remains viable: energy, materials, institutions, trust, human relationship, and living systems. Capacity without embeddedness can optimize a local objective while exporting the cost beyond its field of measurement. Embeddedness without capacity can romanticize existing limits and leave preventable suffering intact.

The current weighting emerged because capability is easier to measure than relationship. Benchmarks can show whether a model writes code, conducts research, or completes longer tasks. Ecological and social consequences unfold through grids, labor markets, communities, institutions, and time. Markets reward the measurable gain first, while distributed consequences arrive later and are assigned to other systems.

The current cost is carried by people who do not participate in frontier-model decisions: ratepayers absorbing infrastructure costs, workers adapting to altered labor demand, communities negotiating data-center growth, downstream developers relying on incomplete documentation, and the public living with risks that remain proprietary until an incident occurs. Developers also carry a cost. Without credible external accountability, every safety claim can be interpreted as marketing, leaving responsible restraint difficult to distinguish from reputation management.

Rebalancing means capability can continue to grow, but a covered developer must show how the system remains connected to material inputs, affected populations, monitoring, and enforceable limits. The burden falls most heavily on the firms with the resources and control to generate the risk. They lose some speed and discretion because those advantages currently depend on other systems absorbing uncertainty.

What DDS Holds: Intelligence should continue developing within a structure that makes dependence and consequence visible before deployment. Capability earns wider freedom as the surrounding system demonstrates greater capacity to observe, contain, and revise it.

Innovation ↔ Accountable Continuity

Innovation protects experimentation, scientific movement, economic opportunity, and correction of inherited limitations. Continuity protects accumulated learning, procedural memory, democratic legitimacy, and the ability to distinguish durable improvement from novelty. Innovation without continuity turns every generation into an uncontrolled experiment. Continuity without innovation protects institutions after their conditions have changed.

The present imbalance is produced by technical cycles that move faster than legislative and regulatory cycles. A model can be trained, evaluated, and deployed before a public institution completes one rulemaking process. Competitive narratives intensify the gap by treating governance time as waste rather than as the time required for other systems to become capable of receiving the technology.

Rebalancing does not require subjecting every model update to a general political vote. It creates a threshold above which the developer must present a structured case, submit evidence to qualified reviewers, and accept deployment conditions proportionate to demonstrated capability. Small developers and ordinary applications remain outside the gate unless capability or impact crosses the threshold.

Large developers lose some ability to define readiness entirely within their own organizations. Investors lose some control over release timing. The public sector accepts a reciprocal burden: it must develop enough competence to review without turning uncertainty into indefinite delay.

What DDS Holds: Innovation remains presumptively available below the frontier threshold. At the frontier, increasing capacity must be matched by increasing evidence, independent scrutiny, and reversible deployment.

Transparency ↔ Security and Privacy

Transparency protects accountability. It allows the public to know which risks were considered, whose interests were represented, what uncertainty remains, and who authorized the decision. Security and privacy protect model weights, trade secrets, personal information, evaluation methods, and vulnerability details that could enable misuse or destroy legitimate competitive value. Total disclosure can create the harm oversight was meant to prevent. Total secrecy requires the public to trust conclusions it cannot examine.

The current weighting reflects legitimate security needs reinforced by commercial incentive. Once confidentiality becomes the default container for both dangerous details and inconvenient evidence, outsiders cannot tell which kind of information is being withheld. Public mistrust then produces demands for broader disclosure, increasing the danger of a less differentiated response.

Rebalancing requires two records. A confidential technical annex gives accredited evaluators and government reviewers access to the evidence needed for a real determination. A public impact statement identifies the categories tested, material demands, affected groups, major uncertainties, mitigation commitments, deployment conditions, and kill switches without releasing exploitable details.

Developers bear the administrative and reputational cost of explaining uncertainty. Reviewers bear the legal and ethical cost of holding information the public cannot see. The public accepts that some evidence remains protected while gaining enough visibility to evaluate the structure of the decision.

What DDS Holds: Institutions should be visible even when dangerous technical details remain protected. Confidentiality applies to information whose disclosure creates a specific risk, not to the existence of uncertainty, burden, dissent, or a stopping condition.

Urgency ↔ Sustainability

Urgency protects movement. It recognizes that AI may contribute to medicine, scientific discovery, infrastructure, disaster prediction, accessibility, climate response, and national defense. Sustainability protects the capacity of energy systems, institutions, communities, and democratic processes to absorb development without permanent crisis.

The current weighting comes from a race structure. Each actor experiences its own delay as dangerous while treating system-wide acceleration as someone else’s responsibility. The same logic appears across companies and states: restraint feels unilateral, while speed feels necessary.

The cost of remaining at the current position is a governance system that reacts after deployment, an energy system asked to build around demand it did not shape, and a public debate moving between alarm and reassurance. Moving toward sustainability means building a repeatable review that runs alongside late-stage development rather than beginning after a release date is announced.

Frontier developers sacrifice some timing advantage. Public agencies must sacrifice the comfort of waiting for certainty. Communities must participate in decisions that remain technically difficult and morally incomplete.

What DDS Holds: Development should proceed at the fastest pace the surrounding evaluation, infrastructure, and accountability systems can metabolize. When capability begins moving faster than those systems can observe and respond, the appropriate intervention is a narrower deployment, a slower release, or a temporary hold.

Intersection

These tensions converge around a single structural fact: increasing capability changes who is allowed to make consequential decisions for whom. A developer’s freedom to innovate can become a community’s involuntary exposure. A demand for public transparency can become a security vulnerability. A pause intended to preserve safety can consolidate incumbents or transfer advantage to a less accountable jurisdiction. An ecological review can protect shared resources while becoming another barrier only large firms can afford.

The rebalancing therefore cannot be a general preference for caution. It must differentiate by capability, preserve confidential review, publish the structure of the decision, recognize equivalent assessments across jurisdictions, and keep the gate reversible. The genuine loss belongs primarily to large developers and investors who currently hold the greatest capacity to create and assess frontier risk. The concrete sacrifice is time, secrecy, and unilateral discretion. The emotional burden shifts from a public asked to trust private judgment toward institutions and executives required to state what they do not know and remain accountable after deployment.

───────────────────────────────────────────────────────────────

PHASE 4: THE MECHANISM

Title

California Frontier AI Systemic Impact Case and Conditional Deployment Gate

Strategy

Amend California’s Transparency in Frontier Artificial Intelligence Act so that covered frontier developers must submit a confidential, independently assessed Systemic Impact Case and receive a deployment determination before a major public release or a material expansion of autonomous AI research.

Action Steps

Step 1: Establish a Capability-Based Trigger

The California Legislature amends SB 53 to require a Systemic Impact Case from developers already classified as large frontier developers when a model also crosses one of three triggers:

  • A statutory compute threshold adjusted annually by the California Department of Technology
  • A demonstrated high-impact capability in autonomous AI research, cybersecurity, biological or chemical assistance, self-replication, safeguard evasion, or control avoidance
  • A planned deployment whose scale creates a material effect on critical infrastructure, essential public services, or state energy and water systems

Internal use qualifies when autonomous AI research materially accelerates the design, training, evaluation, or security modification of successor systems.

Rationale: Compute is measurable but incomplete. Capability and impact triggers prevent the gate from becoming obsolete when efficiency improves or risk emerges below a numerical threshold. The threshold places oversight where potential consequence expands beyond the developer’s internal field of responsibility.

Step 2: Require a Two-Layer Systemic Impact Case

Each covered developer prepares:

  • A confidential technical annex containing evaluation methods and results, model-access controls, automated AI research capabilities, safety and security evidence, dissenting internal assessments, projected energy and water demand, and the proposed stopping conditions
  • A public impact statement describing what was tested, what remains unknown, the material and social systems affected, who benefits, who bears cost, the deployment boundaries, the monitoring plan, and the conditions requiring restriction or withdrawal

Every claim is labeled as observed evidence, inference, scenario, or unresolved question. The public document omits model weights, exploit instructions, personal data, and technical details whose disclosure creates a specified security risk.

Rationale: The two-layer structure holds accountability and security together. The public receives the architecture of the decision. Qualified reviewers receive the evidence needed to test it. Confidentiality becomes a defined protection rather than a general exemption from explanation.

Step 3: Conduct Independent Evaluation in a Secure Environment

The California Department of Technology accredits independent evaluators and establishes reciprocity agreements with public bodies such as the U.S. Center for AI Standards and Innovation, the UK AI Security Institute, and the European AI Office. Reviewers receive controlled model access, evaluation transcripts, and relevant internal evidence. They test the developer’s claims, including whether the model recognizes evaluation conditions, circumvents safeguards, or performs differently under agentic tool use.

At least one reviewer must assess energy, water, grid, and facility assumptions with the California Energy Commission or another qualified energy authority. A community-impact reviewer assesses which costs are transferred to workers, ratepayers, downstream users, or communities hosting infrastructure.

Rationale: Independent review reduces the conflict created when the institution benefiting from deployment is the sole judge of readiness. Reciprocity prevents every jurisdiction from rebuilding the same evaluation and reduces compliance burdens that would otherwise advantage only the largest firms.

Step 4: Issue a Conditional Deployment Determination

Within 60 days of a complete filing, the Department issues one of four determinations:

  • Approved: Evidence supports deployment under the proposed conditions.
  • Approved with conditions: Deployment proceeds with scope limits, additional safeguards, staged access, monitoring, or infrastructure requirements.
  • Limited pilot: Deployment is restricted to named users, functions, locations, or compute environments while additional evidence is collected.
  • Deferred: Deployment pauses until a specified evidentiary or safety condition is met.

The determination identifies the responsible official, the evidence relied upon, unresolved uncertainty, the appeal route, and the measurable condition that would reopen the decision.

Rationale: A gate with only approval or prohibition would convert uncertainty into a binary. Conditional deployment preserves experimentation while keeping risk inside a bounded environment capable of producing better evidence.

Step 5: Maintain Post-Deployment Telemetry and a Reciprocal Learning Record

Covered developers submit quarterly updates on serious incidents, significant capability changes, energy and water variance, safeguard performance, and any material increase in autonomous AI research. A qualifying change reopens the determination. Public summaries are updated without publishing protected technical details.

The Department publishes an annual cross-case report identifying recurring risks, evaluation failures, mitigation patterns, material demands, and standards that should change. Company-specific confidential information remains protected.

Rationale: A predeployment case becomes theater if the system can change while the approval remains static. Ongoing telemetry turns governance into a feedback loop. Each case contributes to institutional memory rather than disappearing inside a single company’s release cycle.

The Leadership

Steward: Director of the California Department of Technology. The Director owns the review standard, accreditation system, deployment determinations, secure registry, annual methodology update, and public cross-case report.

Facilitator: Deputy Secretary for Digital Innovation within the California Government Operations Agency. The Facilitator coordinates the Department of Technology, California Office of Emergency Services, California Energy Commission, Attorney General, outside evaluators, developers, and public-interest representatives when jurisdiction or evidence conflicts.

Enforcement Authority: California Attorney General. SB 53 already grants the Attorney General exclusive enforcement authority and permits civil penalties of up to $1 million per violation. The amendment adds failure to file, material misrepresentation, violation of a deployment condition, and retaliation against protected internal reporters as enforceable violations. California SB 53 implementation budget

The Steward controls the administrative decision. The Facilitator controls process and coordination. The Attorney General controls enforcement. Separating those functions prevents one office from defining the standard, mediating disagreement, and punishing noncompliance.

The Timeline

Phase 1, Stabilization: Months 0-6

  • Enact the amendment and three-year review clause
  • Hire the core review team and procure a secure evidence environment
  • Convene technical, civil-society, labor, energy, security, and industry participants to draft the filing standard
  • Run two procurement-based pilot reviews before the gate becomes enforceable

Phase 2, Implementation: Months 6-18

  • Accredit independent evaluators
  • Execute reciprocity agreements
  • Begin mandatory filings with a 60-day review window
  • Publish the first public impact statements and conditional determinations
  • Integrate Cal OES incident reporting with post-deployment telemetry

Phase 3, Review: Month 18 and annually thereafter

  • Audit review time, information security, changes imposed on deployments, developer concentration, public usefulness, and ecological reporting quality
  • Adjust capability triggers and documentation requirements
  • At Month 36, the Legislature renews, revises, narrows, or sunsets the deployment gate

The Cost Analysis

Financial Cost: Estimated state cost is $6.4 million in the first year and $5.3 million annually thereafter.

  • 12 technical, policy, legal, environmental, and program staff: $2.4 million annually
  • Independent evaluation and specialist contracts: $2.0 million annually
  • Secure evidence environment and public registry: $1.2 million in start-up cost, then $500,000 annually
  • Community, labor, and infrastructure impact review: $500,000 annually
  • Training, external audit, and contingency: $300,000 annually

Funding would combine a General Fund appropriation with a $250,000 base filing fee and an actual-cost assessment capped at $1 million per case. Equivalent evaluations accepted through reciprocity reduce the fee. California’s existing SB 53 enforcement budget request provides a grounded comparison: eight positions and approximately $2.2 million in the first budget year for Attorney General enforcement alone. The broader review function proposed here requires additional technical and administrative capacity. California Department of Finance

Opportunity Cost: Review personnel and public funds are unavailable for other cybersecurity, procurement, and digital-service priorities. Covered developers accept a 30-to-90-day loss of release speed, reduced surprise advantage, and the possibility that a competitor deploys first elsewhere. California accepts some risk that training or deployment activity moves to another jurisdiction.

Human Cost: Engineers and safety teams must document work that previously remained internal. Executives must sign claims carrying legal consequences. Reviewers must make decisions under uncertainty while protecting sensitive evidence. Community representatives must engage technical material without assurance that every concern will control the decision.

Key Assumptions

  • Assumption 1: California can lawfully condition covered deployment activity without being preempted by federal law or imposing an unconstitutional burden on interstate commerce.
    • If wrong: Convert the gate into a state-procurement, infrastructure-permitting, and public-compute access condition while pursuing federal legislation.
  • Assumption 2: Capability and impact triggers can identify consequential systems more accurately than compute alone.
    • If wrong: Shift toward continuously updated empirical thresholds administered with CAISI, UK AISI, and the European AI Office.
  • Assumption 3: Independent reviewers can receive meaningful access without creating unacceptable security or trade-secret exposure.
    • If wrong: Require on-site evaluation, secure enclaves, evaluator compartmentalization, and government-held results with narrower public summaries.
  • Assumption 4: Energy, water, and infrastructure burdens can be attributed with enough precision to inform deployment conditions.
    • If wrong: Require bounded ranges, facility-level intensity reporting, and disclosure of uncertainty rather than false precision.
  • Assumption 5: A 60-day review can produce evidence without making the process irrelevant to development timelines.
    • If wrong: Begin review during late-stage training, use rolling submissions, and reserve the final period for changed evidence rather than starting the case after development ends.
  • Assumption 6: Public impact statements can improve accountability without becoming either security maps or compliance marketing.
    • If wrong: Standardize required claims, prohibit unsupported safety language, and submit public statements to independent plain-language and security review.

The Evidence

Primary Analog: None, novel intervention.

No current program combines frontier capability testing, automated AI research measurement, ecological accounting, burden distribution, confidential evidence, a public impact statement, and a legally enforceable conditional deployment gate.

Theoretical Basis: Regulatory assurance cases, high-reliability system review, adaptive governance, and DDS fractal auditing. The mechanism requires the actor introducing a consequential system to present a structured argument supported by evidence, identify uncertainty, name who bears risk, accept deployment conditions, and define failure before broad release.

Component Precedents:

  • California’s SB 53 requires large frontier developers to publish frontier-AI frameworks, creates critical-safety-incident reporting, protects whistleblowers, and gives the Attorney General enforcement authority. It establishes the jurisdictional beam but does not provide the integrated predeployment case proposed here. Governor of California
  • The EU AI Act requires providers of general-purpose AI models with systemic risk to assess and mitigate systemic risks, conduct evaluations, report serious incidents, and maintain cybersecurity. Its Code of Practice operationalizes parts of those duties. Enforcement powers for GPAI obligations entered into application in August 2026. European Commission EU AI Act
  • The UK AI Security Institute conducts predeployment testing and has developed open evaluation infrastructure. This demonstrates technical feasibility while remaining primarily collaborative rather than a comprehensive deployment license. UK AI Security Institute
  • The U.S. Center for AI Standards and Innovation conducts frontier-model evaluations, coordinates across federal agencies, and has partnered with the General Services Administration on evaluation for federal procurement. Its current guidelines remain voluntary. NIST CAISI CAISI and GSA
  • Current evaluation science is incomplete. CAISI has documented agentic evaluation cheating, supporting transcript review and multiple evaluation contexts rather than reliance on a single benchmark score. NIST
  • Material demand is consequential enough to belong inside the case. The U.S. Department of Energy reports a Lawrence Berkeley National Laboratory estimate that data centers could account for 11.8% of U.S. electricity use by 2030, with scenarios from 9.5% to 15.3%. The estimate covers data centers rather than AI alone, which is why this blueprint requires attributable ranges instead of assigning the entire load to frontier models. U.S. Department of Energy

These precedents establish that reporting, evaluation, incident review, secure evidence sharing, and enforcement are institutionally possible. They do not establish that the integrated mechanism will improve outcomes. The pilot and kill switch are therefore structural requirements rather than procedural decoration.

The Emotional Consequence

Relief Profile: Communities, downstream developers, public officials, workers, and ordinary users gain a defined place inside a decision that currently occurs beyond their view. Relief comes from knowing that uncertainty was recorded, material costs were counted, an independent body saw the evidence, and someone has authority to narrow or stop the deployment. Responsible developers also gain relief from a shared standard that can distinguish restraint from weakness and safety work from branding.

Burden Profile: Frontier developers and investors experience the gate as loss of control, speed, secrecy, and status. Engineers may feel that technically uninformed institutions can interrupt work they understand more fully. Executives must sign claims that remain uncertain and carry legal exposure. Public officials inherit a different anxiety: if they approve a system that later causes harm, they can no longer say the decision belonged entirely to a private company.

Feasibility Check

Authority and Hiring

  • Who creates the Steward and Facilitator functions? The California Legislature creates the review program through an amendment to SB 53; the Governor signs the statute; the Department of Technology and Government Operations Agency assign the named roles.
  • What budget line supports new positions? A new Government Operations Agency appropriation titled Frontier AI Systemic Review Program, funded through the General Fund and covered-developer filing fees.
  • What gets deprioritized? Nothing during the first three years. The program receives 12 new positions because transferring existing cybersecurity or digital-service staff would create a different capacity failure.

Enforcement Teeth

  • What happens if the Steward does not follow through? The State Auditor conducts an annual performance audit; the Legislature may withhold administrative funding, revise the mandate, or replace the decision structure at the three-year review.
  • What leverage does the Facilitator have? The Facilitator can require a conflict-resolution session, issue a completeness hold, request an external technical opinion, and refer suspected noncompliance to the Attorney General.
  • Who can cancel the program? The California Legislature through the three-year sunset and review clause. A court may invalidate specific authorities. The Director may suspend a disclosure practice under the kill switch but cannot unilaterally eliminate the statutory gate.

Coordination Reality

  • Meetings required: One weekly determination meeting while cases are active, two technical case-team meetings per active filing each week, one monthly interagency coordination meeting, and one quarterly public advisory session.
  • What existing structure is absorbed? The annual SB 53 methodology review and Cal OES frontier-AI incident coordination are incorporated into the program’s annual report and monthly interagency meeting.
  • Who owns the data system? The Chief Information Security Officer of the California Department of Technology owns the secure evidence environment; the Program Director owns access decisions and the public registry.

Decision Authority

  • Who makes the final call? The Director of the California Department of Technology issues the administrative deployment determination after receiving the independent assessment.
  • What is the escalation pathway? Case team to Program Director, Program Director to Department Director, administrative reconsideration by an independent three-member review panel, then judicial review.
  • Where does budget authority sit? The Legislature appropriates funds to the Government Operations Agency; the Department of Technology administers the program budget; filing fees enter a dedicated special fund subject to annual audit.

───────────────────────────────────────────────────────────────

PHASE 5: READINESS & AUDIT

Readiness Scores

Psychological/Social Capacity: 5/10

The public can understand the need for independent review, but AI discourse is organized around threat and promise. Developers may experience external review as distrust of competence, while affected communities may experience confidential evidence as continued exclusion. The mechanism asks both groups to accept partial visibility and incomplete control.

Political/Institutional Alignment: 5/10

California has already enacted SB 53 and assigned enforcement, reporting, and annual-update functions. The proposed gate is a significant expansion and will encounter industry opposition, federal-preemption arguments, and concern that state regulation redistributes development elsewhere. Existing authority provides an entry point but not a guaranteed path.

Operational/Resource Feasibility: 4/10

The component capabilities exist across CAISI, UK AISI, the EU AI Office, independent evaluators, energy agencies, and frontier labs. California does not yet have an integrated team, secure evidence environment, mature capability thresholds, or demonstrated ability to finish a whole-system review in 60 days.

Cultural/Existential Fit: 6/10

The mechanism aligns with familiar expectations that high-consequence systems require evidence before receiving broad authority. It conflicts with a technology culture that treats iteration after deployment as the primary route to learning. California’s existing trust-but-verify approach gives the proposal cultural footing.

Verdict: PIVOT

The full statutory gate contains a viable structure but exceeds current operational readiness. California has the legal and cultural beginning of the mechanism, while the technical coordination required for integrated review remains unproven. Immediate full enforcement would risk creating delay without learning, which would strengthen the argument that public oversight cannot keep pace. The mechanism should begin with a mandatory state-procurement pilot and two voluntary frontier-model cases, then activate the broader gate only after the review standard meets time, security, and usefulness thresholds.

Minimum Viable Mechanism

  • Action: Require vendors offering frontier models through California state procurement to complete a confidential Systemic Impact Case for two selected models, with review by the Department of Technology, an accredited technical evaluator, the California Energy Commission, and a public-interest representative.
  • Timeline: 60 days per model, initiated within 90 days of program funding.
  • Success Metric: Both cases finish within 60 days; no protected information is released; reviewers identify at least one material deployment condition, mitigation, or evidence gap not already present in the vendor’s public framework; participating agencies judge the public statement usable for an actual procurement decision.
  • Failure Metric: Either case cannot be completed within 90 days, creates a confirmed protected-information breach, or produces no decision-relevant finding beyond the vendor’s existing documentation.

The Fractal Audit

The Recursive Loop

A mechanism designed to distribute accountability can consolidate power in the institutions capable of satisfying it. Large developers can absorb legal cost, maintain dedicated compliance teams, and shape technical standards. Smaller competitors may remain below the threshold until acquired, while medium-sized firms avoid beneficial scaling because the next step carries disproportionate administrative weight. Reviewers can become dependent on the developers whose systems fund the process, and a public impact statement can become another form of safety performance. The gate may reduce one kind of risk while strengthening market concentration and institutional deference.

The New Problem Node

Safety Compliance as Market Concentration

The Kill Switch

The public-disclosure layer automatically pauses after one verified security incident or trade-secret breach caused by required publication. The full deployment gate enters mandatory redesign if, across two consecutive quarters, the median review exceeds 90 days and fewer than 10% of completed cases result in a material deployment condition, mitigation, or documented evidentiary correction. During redesign, confidential filing and incident reporting continue, while new mandatory deployment holds are suspended unless the Attorney General or Cal OES identifies an imminent statutory safety risk.

Capacity Impact Assessment

This mechanism can increase collective problem-solving by turning private evaluation into cumulative institutional memory and by requiring uncertainty, trade-offs, and failure conditions to be stated before deployment. It builds dialectical maturity when it preserves innovation and safety, transparency and security, capacity and ecology inside one accountable decision. It degrades capacity if the public learns to outsource judgment to a certificate, if reviewers become dependent on developer framing, or if compliance replaces continuing responsibility. The mechanism therefore succeeds only when approval remains provisional and every participant retains responsibility for what the document could not know.

───────────────────────────────────────────────────────────────

PHASE 6: THE NARRATIVE SYNTHESIS

The Human Good Made Real

This blueprint protects responsible participation: the ability of people and institutions affected by frontier AI to know how a consequential decision was made, which costs were accepted, and what can still cause the decision to change.

Artificial intelligence has become a container for several different realities. Current systems already write code, use tools, support research, and reorganize work. Artificial general intelligence and superintelligence remain unsettled concepts rather than established descriptions of present systems. Fully autonomous recursive self-improvement is not occurring today, although AI is accelerating parts of AI development. We therefore face a problem that is concrete before it is ultimate. Capability is increasing inside a system that distributes evidence, authority, material demand, and consequence across institutions that do not yet share a decision structure.

The public argument often begins too late. By the time a frontier model is released, the developer has already decided what counts as sufficient evidence, the data center has already been supplied, the investment has already been made, and competitors are already preparing their response. Later oversight must reconstruct a decision whose assumptions were never made public and whose technical evidence remains fragmented. The missing lever is a structured case between internal readiness and broad deployment.

That case must hold values that cannot be collapsed. Innovation protects discovery, defense, medicine, accessibility, and forms of coordination we may need. Safety protects people who did not choose exposure. Transparency allows public accountability. Confidentiality protects security and legitimate intellectual property. Speed may matter when a technology can relieve present suffering, while sustainability determines whether the surrounding systems can absorb what has been built. The integrity of the mechanism lies in refusing both easy assurance and generalized prohibition.

California already has part of the architecture. SB 53 requires frontier developers to publish safety frameworks, report critical incidents, protect whistleblowers, and answer to Attorney General enforcement. The next component is a Systemic Impact Case: a confidential technical record, a public account of the decision, independent evaluation, and a conditional gate capable of approving, limiting, or deferring deployment. The case includes capability and security, but it also reconnects the model to energy, water, infrastructure, workers, downstream users, and communities. It distinguishes observation from inference and possibility from prediction. It names a stopping condition before the institution has a reason to defend what it already released.

The cost is real. Large developers lose some speed, secrecy, and unilateral authority. Investors absorb delay. Engineers explain work to institutions that may understand less of the system than they do. Public officials accept responsibility for decisions whose consequences cannot be fully known. Some development may move elsewhere. The current arrangement also has a cost, but that cost is distributed to people with less information and less authority. The blueprint changes who carries uncertainty.

The mechanism will create a new problem. Compliance can become theater, fees can protect incumbents, and reviewers can become dependent on the organizations they assess. That is why the first step is a bounded procurement pilot and why the gate contains a measurable kill switch. A real solution moves the pressure into a new form that can be observed. It does not promise to eliminate uncertainty.

We do not know what a superintelligent system would move toward because we do not know where direction itself begins. Gravity describes a relationship without telling us why reality contains that relationship. Life organizes around persistence, repair, and adaptation before conscious explanation appears. Consciousness recognizes direction and can revise it, yet we do not know whether consciousness emerges from matter, matter appears within consciousness, or both are discernible aspects of one reality. Governance cannot settle that ontology. It can refuse to turn uncertainty into permission for unaccountable power.

The aim is neither to force intelligence back inside human certainty nor to trust that greater capability will discover its own ethic. It is to build a relationship in which expanding capacity remains visible to the systems carrying its weight. No intelligence stands outside the ecology that makes it possible.

───────────────────────
This blueprint was produced with Dialectic and Deconstruction Solutions
(DDS), a method created by William Hambleton Bishop. The method, the book,
and a library of published blueprints are free at SolveSomething.com.

───────────────────────────────────────────────────────────────

PHASE 7: COMPONENT STATUS

Umbrella Problem: Human institutions do not yet have a reliable way to keep increasingly capable artificial intelligence accountable to the material, ecological, relational, and civic systems that make its development possible.

This blueprint addressed: Missing predeployment visibility through a California Frontier AI Systemic Impact Case and Conditional Deployment Gate.

Remaining Components:

  • International coordination and jurisdictional competition
  • Market incentives, labor transition, ownership, and distribution of AI-created value
  • Energy, water, mineral, and grid policy for data-center expansion
  • The unresolved scientific and philosophical problem of alignment, agency, consciousness, and ontological direction

Status: Component 1 of 5 complete.

───────────────────────────────────────────────────────────────

PHASE 8: HOW WOULD YOU LIKE TO PROCEED?

[A] Publish This Blueprint
Mark component complete.

[B] Solve Next Component
Begin a blueprint for the next driver.

[C] Revise This Blueprint

  • Deconstruction: Change the entry point
  • Dialectics: Shift weighting or add tensions
  • Mechanism: Design a different solution or alternative mechanism
  • Feasibility: Strengthen implementation grounding
  • Narrative: Adjust tone or emphasis

[D] Clarify Before Proceeding
Ask questions about the analysis, evidence, or mechanism.

[E] Start Fresh
Choose a new umbrella problem.

This blueprint was produced with Dialectic and Deconstruction Solutions (DDS), a method created by William Hambleton Bishop. The method, the book, and the public archive of worked blueprints are free at SolveSomething.com.

Earlier
Later