Changes confirmed medium confidence

Meta Extends Scaling Framework to Training and Open-Weight Releases

The updated policy adds containment controls before high-risk training runs and expands how Meta says it will assess models whose weights can be modified after release.

Edited by Tyronne Panaino

Meta updated its Superintelligence Scaling Framework on October 2 to extend loss-of-control precautions into model training and evaluation and to expand its treatment of open-weight releases. The policy change matters to researchers, safety teams and organizations assessing Meta models because it distinguishes controls applied before a high-risk training run from the separate risks created when model weights can be modified after release.

The official Meta Research update says the framework sets capability thresholds and safety or security requirements before training begins and before deployment. Meta connects the revision to commitments made at the White House that week, but the new page is a company-specific implementation statement rather than another account of the wider multi-company accord.

Containment now begins before high-risk training

Before a reinforcement-learning training run, Meta says it will determine whether the model being trained could compromise its environment. When that risk is present, the company says the run must use mitigations before it starts.

The listed controls include internally approved sandboxes that are vetted for vulnerabilities, real-time secure logging of model outputs and automated monitoring across high-risk rollouts. Meta says those monitors can halt a run when they detect severe misbehavior or an attempted escape from containment. It also says the monitoring systems will be stress-tested to assess whether they can detect unauthorized model actions.

This moves the stated control point earlier than public deployment. A model can interact with tools, graders and surrounding infrastructure during reinforcement learning or evaluation, so a policy that only evaluates the final release would leave an earlier operational stage outside the framework. Meta's update says its loss-of-control requirements now cover that stage as well.

Open weights receive a deployment-specific assessment

The revised framework also expands Meta's discussion of what can happen after model weights are released. The company identifies several ways a weight holder could weaken refusal behavior, including resampling, prefilling outputs and additional fine-tuning. It says risk assessment should therefore account for the model's specific deployment context rather than assuming the original refusal behavior remains fixed.

Meta also says it uses outside experts for threat modeling where appropriate and consults government bodies and other stakeholders. For biological and chemical risk, the company says it considers both the model's capabilities and how a deployment could contribute to proliferation. These are stated assessment practices, not public evidence that every future release has passed an independent review.

Board oversight is still a future step

The October update separates changes already described in the framework from governance work planned over the coming months. Meta says it will establish a new board-level AI committee to review future framework changes and check whether operations conform to the company's standards. The source does not say that this committee already exists or provide a completion date.

That timing is important. The revised training and open-weight provisions are the current documented policy, while the committee is a commitment for later implementation. Readers should not collapse those two statuses into a claim that the announced oversight structure is already operating.

Evidence quality and limitations

The evidence is a single actor-controlled Meta publication. It establishes what the company says its framework now requires, but it does not independently verify sandbox quality, logging completeness, monitoring effectiveness, auditor access or consistent implementation across training runs. It also does not report incidents, pass rates or external evaluation outcomes for the new controls.

The next useful checkpoints are publication of the revised framework details, evidence that the promised board committee has been established and independent evaluation of the containment and monitoring process. Those checks would show whether the policy language is functioning as an operational control rather than remaining only a documented commitment.

Status

Confirmed. Meta has published the framework update and described the new training and open-weight provisions. Internal confidence is medium because no independent implementation or effectiveness evidence was fetched.

Sources

Update note: Last reviewed October 8, 2026. We will revise this post when Meta publishes evidence of implementation, external review or the planned board committee.

Sources

Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.

More Changes coverage