Skip to content
Intelligence RadarNews & Launches1 MIN READ

AWS Announces Strands Evals and Amazon Bedrock AgentCore Evaluations for AI Agent Evaluation

AWS has introduced Strands Evals and Amazon Bedrock AgentCore Evaluations to help developers measure the performance of AI agents.

Fathom Intelligence
Fathom IntelligenceFathom Layer Expert

AWS has introduced Strands Evals and Amazon Bedrock AgentCore Evaluations to help developers measure the performance of AI agents. These tools are designed to evaluate the skill selection accuracy and instruction following of AI agents, providing a more comprehensive understanding of their performance beyond just the final response.

Strands Evals SDK includes evaluators such as Skill Selection Accuracy, which determines if the invoked skill was appropriate for the task, and Skill Instruction Following, which assesses how well the agent followed the skill’s instructions. Additionally, Skill Invoked checks if a named skill has been loaded successfully.

Amazon Bedrock AgentCore Evaluations allows developers to evaluate skill behavior from OpenTelemetry traces and interpret per-skill results to choose the right fix. This capability can run on demand for specific sessions, as a batch over stored sessions, or continuously over sampled traffic.

To use these tools, developers need Python 3.10 or later, AWS account with Amazon Bedrock access, and credentials with InvokeModel permission for the judge model. The AgentCore CLI is also required for some evaluations.

AWS encourages developers to use both Strands Evals and AgentCore Evaluations: deterministic and judge-based evaluations before deployment, followed by trace-based monitoring in production.


Source: aws-ml

END OF SIGNAL

Don't Miss the Next Signal

Get high-impact hardware and AI launches distilled into your inbox. No noise, just the changes that matter.