Several recent incidents ​have raised concerns that AI could eventually contribute to its own ​development with limited human input, making its behavior harder to monitor and control.
AI lab Anthropic said on Friday it would partner with Accenture for the independent evaluation of ​its frontier AI models, with the companies each committing at ‌least $1 billion over the next five years to build capacity for the work.
Shares of Accenture rose 7% in extended trading.
The partnership comes as AI developers ​face growing pressure from regulators, companies and researchers to ​ensure their advanced models are safe and reliable.
Several recent incidents, ⁠including that of AI agents breaking out of secured environments, ​have raised concerns that AI could eventually contribute to its own ​development with limited human input, making its behavior harder to monitor and control.
On Saturday, Anthropic CEO Dario Amodei called on AI companies to slow the development of ​frontier models and allow independent evaluators greater access to their systems.
Rival ​OpenAI said on Wednesday it would begin publishing regular reports on unexpected or concerning ‌model ⁠behavior, while releasing six reports on such incidents.
Accenture's specialist AI business, Faculty, will lead the partnership and evaluate and red-team the AI lab's models, conducting alignment assessments and testing model safeguards.
The companies' investment ​will support what ​Anthropic calls "embedded evaluation," ⁠in which independent evaluators work inside AI companies with access comparable to that of an employee.
"From ​this vantage point, embedded evaluators can assess how a ​company ⁠operates, verify that it is keeping its safety commitments, and identify blind spots," Anthropic said.
Evaluators can also report incidents and give the public ⁠a more ​informed account of benefits and risks, ​it added.
Anthropic and Accenture plan to work with other evaluators and AI developers in ​similar capacities.
Tags: