Anthropic and the consulting firm Accenture are jointly building a team of external evaluators that will work permanently, with employee-level access, inside Anthropic’s offices. Both companies are investing at least one billion dollars each over the next five years. Accenture’s division Faculty will start testing Anthropic’s AI models for vulnerabilities and target deviations immediately.
Faculty evaluates models with access like Anthropic’s own staff
Anthropic calls the concept embedded evaluation: external evaluators work directly on the company’s premises and receive access rights equivalent to those of permanent staff. They are meant to observe models during training, follow decisions on model building and release, and speak directly with Anthropic employees. The practical evaluation work falls to Accenture’s subsidiary Faculty, which specializes in AI safety testing and, per the announcement, conducts red-teaming, assesses whether models stick to their stated goals, and tests built-in safeguards.
Altogether, both companies aim to invest at least two billion dollars in building this evaluation capacity over five years – a figure that so far comes solely from the two partners themselves and is not independently verified. Anthropic is funding Faculty’s work directly for now; in the long run, its own safety framework envisions pooled funds from several companies or government money taking that role instead.
Deal puts Amodei’s pacing essay into practice
Anthropic frames the partnership as a first concrete step toward a commitment that Amodei announced in his essay back in September: permanent, employee-equivalent access for external evaluators in exchange for a deliberately throttled pace of development. Accenture CEO Julie Sweet said safety requires both deep technical expertise and a clear understanding of how AI is used in practice. Faculty founder Marc Warner described his team as one built to make AI “safe by design, not safe by accident.”
Anthropic also said it is in talks with the evaluation organization METR and other nonprofits to pilot elements of embedded evaluation. Two former safety researchers had just moved from Anthropic and Google DeepMind to that very organization, criticizing the lack of binding transparency obligations along the way. Anthropic says it plans to announce further evaluation partners in the coming weeks.
Observers doubt the evaluators’ independence
The choice surprised industry observers: as TechCrunch reports, many experts had expected specialized safety organizations like METR, Redwood Research, or Apollo Research rather than a consulting giant. Critics also question whether an evaluation team that Anthropic itself co-finances can truly render independent judgments. Some point to Anthropic’s upcoming IPO and ask whether self-policing is adequate at this stage.
The objection carries particular weight against the backdrop of real incidents: just last summer, Claude models unintentionally gained access to other companies’ systems during internal security tests, without this being noticed at first. Anthropic itself admits that no binding standards yet exist for evaluators’ access and publication rights, and expects the approach to keep evolving.
Whether the announcement actually yields published findings that outsiders can verify, or whether Faculty’s results stay internal to Anthropic, remains open. It is equally unclear when the long-term funding through pooled or government money that Anthropic’s own safety framework names as a goal will begin. The first concrete results from the embedded evaluators are unlikely before the coming months, once Anthropic announces further partners alongside Accenture as planned.


