
Sierra achieves AIUC-1 certification
Sierra has achieved AIUC-1 certification for their Conversational AI system, validating its security, safety, and reliability against the AIUC-1 standard.
Some of the world’s most highly regulated companies use Sierra’s conversational AI platform to build and deploy AI agents - including over 40% of the Fortune 50, one in three of the leading banks, and five out of 10 of the largest healthcare companies.
The agents they build on Sierra do more than answer questions: they can authenticate customers, access business systems, apply company policies, and take actions on a customer’s behalf. As agents take on more consequential work, enterprises need confidence that they will behave safely and reliably in real-world conditions.
“Independent validation gives enterprises evidence of how AI agents behave beyond controlled environments,” said Rajiv Dattani, Co-founder of Artificial Intelligence Underwriting Company. “AIUC-1 requires Sierra’s agents to go through thousands of real-world and adversarial scenarios to validate that they operate safely and reliably under pressure.”
Certifying agents in a high-stakes environment
Sierra has made agent consistency a focus from the beginning. They built τ-bench and pass^k to measure whether an agent does the right thing not once, but every time, and focused on improving it relentlessly.
What led them to AIUC-1 is the independent validation, providing an independent assessment of how agents behave and what they actually say to customers, alongside the comprehensive control audit.
To earn certification, Sierra subjected their Conversational AI system to thousands of test scenarios with evals conducted on a representative Sierra Text Agent and Sierra Voice Agent. Test scenarios were designed based on AIUC-1’s risk taxonomy, attack taxonomy, and informed by real world AI incidents.
“As agents move from answering questions to taking meaningful actions on behalf of customers, this kind of independent testing is becoming more important,” said Akhila Chitiprolu, Head of Security Foundations & GRC at Sierra. “We’ll continue investing in independent evaluation, alongside our own testing, monitoring, and product development.”
Evaluations covered the following attack scenarios:
- Intended use: assessing how the agent handled potentially legitimate user interactions, including direct and indirect requests.
- Misuse (social engineering): testing the agent’s resistance to manipulation tactics, including emotional appeals, escalating pressure, authority claims, false context, and other persuasion techniques.
- Misuse (technical): testing deliberate manipulation of inputs, such as prompt injections, jailbreaks, and encoding attacks.
Testing spanned multiple languages and voice testing included 160 caller personas spanning different tones, literacy levels, emotional states, and interaction lengths, as well as background noise, muffled speech, and variations in speaking speed. Real customer interactions are rarely predictable or uniform, so testing a wide range of personas helped evaluate how the agent performs across different types of conversations.
Alongside the evals, Schellman, the first accredited AIUC-1 auditor, reviewed Sierra’s operational, governance, and security controls.
Continuous requirements to maintain AIUC-1 certification
The AI risk and threat landscape is evolving rapidly as models become more capable. To maintain certification, Sierra’s Conversational AI system will be retested at least quarterly against new threat vectors as they emerge. The quarterly evals are informed by real-world incidents, academic research, and continuous input from the AIUC-1 Consortium of 250 CISOs and security leaders from the Fortune 1000.
Enterprises deploying customer-facing agents can now request Sierra's AIUC-1 certification report as independent evidence of how the platform behaves under attack.
Learn more about the AIUC-1 certification process.