SentinelLABS recently examined the gap between cybersecurity benchmark scores and operational usefulness. In the months that followed, they observed those same design gaps in the flood of agentic systems hitting the market. This fail-fast development trend treats evaluation as a process that emerges after the proof of concept. In other words, teams build the system, demonstrate that it can work, and defer the harder task of establishing how reliably it performs and where it fails for a later date.[1]
This makes sense in the new world of rapid prototyping and AI experimentation. However, these systems are often brittle or work only until something in the environment changes or they fail to generalize. This raises the urgent question, "How do I fix or improve this one behavior?" inside of a monolithic, non-deterministic pipeline.
For a system that will support security decisions, we believe the most important design choice is to make its performance measurable from the start. That means building the agent so that engineers can identify where an investigation goes wrong, test a correction, and verify that the change improves the complete workflow.
Consider a grounded example: an investigation agent that reviews an alert, queries the SIEM, and returns a verdict that matches an experienced analyst's. That looks like a success. On closer inspection, however, the query used the wrong device identifier and retrieved unrelated events. The verdict may be correct, but the records gathered during the investigation do not support the decision.
So, how would the team building that agent discover the mistake? And after fixing it, how would they know the next model or harness update hadn't brought it back?
Start With the Work the Agent Should Complete - When building agents, attention often goes to the visible machinery: which model, which prompts, which tools, and so on. Those matter, but they're also some of the easier parts to modify. The hardest part to retrofit is the ability to answer two questions: Is the system effective? And if not, what exactly do I change to make it better?
The industry has converged on this from several directions. Frontier lab engineering guidance observes that teams get surprisingly far on manual testing and intuition, but once an agent is in production, building without evaluations breaks down. Users report that the agent "feels worse" after a change, engineering scrambles to see whether the claim has any validity, and it becomes difficult to distinguish real regressions from noise or to test changes against realistic scenarios before shipping [2]. For some, the answer is what they call eval-driven development: scoped tests at every stage, before and during implementation, like test-driven development, which writes tests before the code exists [3].
What that guidance usually leaves implicit is that evaluation discipline is an architectural decision. If you want to measure components individually with rigor, the agent must be built from individually exercisable competencies.
Before we choose a model or write a prompt, we need to define the work we expect the agent to do. For an investigation agent, that means reviewing an alert, gathering the evidence needed to understand what happened, and reaching a conclusion that the evidence supports. When it cannot resolve the case, it should give an analyst a clear account of what it found and what still needs investigation.
On Evaluations - "The first principle is that you must not fool yourself, and you are the easiest person to fool." – Richard Feynman. An evaluation measures a model's performance against a given scenario. More formally, every evaluation decomposes into a task (the point of indecision the system must reason through), a scenario (the environment wrapped around that task), and rubrics (how completion gets judged). This may seem simple on the surface, but LLM evaluations are an emerging discipline that every industry is struggling with. As the Pydantic team puts it, "Anyone who claims to know exactly how your evals should be defined can safely be ignored" [4]. A good eval set asks, "What will a team use this for?" and ensures the system gives sufficient answers consistently.
End-to-End Evaluation: Does the System Perform? An end-to-end evaluation runs the whole system on a realistic workload with grades as the final outcome. For an investigation agent, the natural frame is the human baseline it augments. A SOC analyst starts the day with a queue. Over a shift, they work some number of alerts, dispositioning them into buckets (benign, suspicious, escalate, respond), and they end the day with a measurable amount of work left over. One metric that matches is burndown: of everything the system started with, how much still needs human attention? Alongside it sit the metrics that keep this burndown honest, such as disposition accuracy versus expert ground truth, time-to-verdict, stability, consistency, and cost per alert.
End-to-end metrics like these are often seen as the only metrics that matter. They tell a compelling story, show that a system is worth building, and match the narrative users want to hear. They also catch emergent failures such as compounding errors, context loss, and mis-sequenced steps that only show up when every component runs together against realistic inputs [2,6]. However, end-to-end metrics can aggregate away the information engineers need. When overall performance is bad, or degrades after a change, it doesn’t provide context for what went wrong. A regression in burndown could indicate one of many problems; meanwhile, all we can measure is that we suddenly aren't performing. In a monolithic system, the only answer is to start reading traces and hunting for sources of failure.
Competency Evaluation: What Exactly Do We Fix? To tackle that question, we use the competency evaluation (elsewhere called component-level [6]). A competency evaluation isolates one skill and tests it directly, with its own specific inputs, definition of success, and a wide variety of different cases. These evaluations are helpful because they tell you whether a single part of a larger system is functional and capable across many scenarios. The inverse is also true, which is to say that it tells you which exact skills are underperforming.
End-to-End and Competency Evaluations: What Each Reveals
|
End-to-end evaluation |
Competency evaluation |
|
|
Question answered |
Does the system deliver value? |
Which capability is weak, and how? |
|
Unit under test |
The whole system on realistic workloads |
One severable skill in isolation |
|
Example metric |
Alert burndown, disposition accuracy, escalation precision |
Query correctness rate against a specific source's schema |
|
Primary audience |
Leadership and the teams relying on the system |
The engineers improving the system |
|
Detects |
Emergent, compounding, integration failures |
Localized skill deficits, regressions in one capability |
|
Fails to provide |
Any indication of what to change |
Assurance that the assembled whole works |
A mature program runs both continuously, and many teams add additional layers. For example, trajectory or trace evaluation, which grades the path the agent took rather than only the endpoints [2,6]. These evaluations are also critical, but this article specifically serves to spotlight the two defined above.
A Real-World Example - Consider Sentinel’s Purple AI® Agentic Investigation agent. When invoked, the system runs the investigation in the same way an analyst would. It pulls the surrounding evidence for an alert, enriches that evidence with additional data, forms and submits queries against the organization's telemetry, reasons over what comes back, and lands a verdict with the findings to support it. The tempting way to grade that answer is obvious: does the verdict match what an expert would have decided?
Very tempting, and also very dangerous. Suppose an agent agrees with the expert 95% of the time. That sounds strong, but the agent could just be pattern-matching on alert characteristics instead of retrieving evidence and building intermediate conclusions. A correct verdict reached for the wrong reasons is a failure, and a verdict-level score alone doesn’t tell us how we arrived at a conclusion. This is the same blind spot SentinelLABS identified in Part 1 - the unit of evaluation was more often than not the answer to a question, as opposed to the path taken to reach it.
Decomposing the Problem - So, how do you evaluate competencies in this example? Start from the job you're trying to solve and decompose. On inspection, investigation hinges on many competencies. The agent has to form hypotheses worth testing based on alert data and the state from previous steps. It has to translate those hypotheses into correct, efficient queries against each data source, and it has to understand that source's schema and what the fields it's filtering on actually mean. It must interpret result sets, pick the next pivot, and know when the evidence is sufficient to stop.
Each of these is a distinct competency that may be evaluated on its own.
An (extremely abbreviated) tree:
Investigation
├─ Hypothesis generation
├─ SIEM query construction (per-source eval suites)
├─ Investigation efficiency (tool-call budget)
├─ Adherence to organizational policy
├─ Confidence & evidence sufficiency judgment
└─ etc.
Now Sentinel’s evaluations can become concrete. "SIEM query construction" means many cases, each specifying an investigative intent, a target data source, and ground truth for what a correct query returns; we grade it automatically by running the query against representative data and comparing results, and we back it with experts who wrote their own queries for the same asks.
For example, consider the device-ID failure we mentioned above. A competency case could alert the agent with a device ID and ask it to retrieve the related authentication events from a representative SIEM dataset. The case already defines which events a correct query should return.
If the agent queries the wrong identifier, it fails the case. The query may run successfully and return plausible data, but it does not return the expected events. This turns a failure that was previously hidden inside a successful investigation into something the team can measure directly. Once the team fixes the problem, the case stays in the evaluation suite to ensure future system changes are consistently measured.
Even just a few dozen well-chosen evaluations per capability, drawn from real failures, can provide a strong initial signal. Each competency gets its own suite since we know models may struggle in these areas.
Where to Invest Evaluation Effort First - Not every node in the tree deserves equal attention. When setting up your own evaluations, we suggest asking two questions:
Is this actually hard for models? Some competencies are commodities today. Summarization is the canonical example. Modern LLMs summarize nearly anything competently, and the difference between a good and a great summary rarely moves outcomes. Contrast that with SIEM query construction, a real technical hurdle, with objective failure (the query is wrong, slow, or returns the wrong rows), high downstream impact (a bad query corrupts the whole investigation), and known model weakness on dialect and schema-specific syntax. Similar challenges belong in your evaluation cases.
Can you define a defensible standard? Some competencies rely on direct checks against known facts, while others require expert judgment about whether an agent’s decision was justified by the evidence available. Evaluation effort pays off where success is decidable. Query construction has crisp ground truth (run it, compare results). Evidence-sufficiency judgment is softer, but experts mostly agree on clear cases, so a rubric plus expert-labeled examples works. Where even experts can't agree what "good" means, an eval suite just encodes noise. Before evaluating, you must first clearly define success [10,11].
A useful heuristic is to simply combine the two and prioritize competencies by impact of failure × probability of model failure × decidability of ground truth. Wherever possible, grade binary pass/fail per case instead of arbitrary quality scores, since binary judgments are easier for experts to make consistently, easier to trend, and harder to argue with [10]. If your system is already in the wild, you can still build cases - but analysts recommend you prioritize observed failures over successes [2,3].
The Value of Domain Expertise - Every competency evaluation needs to answer the question "what does good look like?" The person who can say whether a generated query succeeded (not just executed, but "all things considered, this is equal to or better than what I would have written in my day job") is the person who writes those queries well.
This is why many companies are spending heavily on contracted expert annotators. They're buying, at market rates and arm's length, the domain judgment that their evaluations require. In reality, most AI engineers are not domain experts, and many teams building agents don't have these resources in abundance. This is where being a company of cyber professionals is our advantage: for nearly any skill a security agent needs, someone here already performs it at an expert level daily. That expertise is the standard every one of our evaluations is built against.
When you identify a competency, identify its experts, understand how the work actually functions, and co-author cases with them. Evaluations built this way stay true to real life, improving generalization when they reach the field, because they're tested against reality rather than the builder's imagination of what the field might look like.
Prove It Before Production - A well-built competency tree carries a risk because it makes the evaluation effort feel finished. However, evaluation is a constant game of discovering scenarios, hill climbing performance, and monitoring for regressions. Inside SentinelOne®, once an agent has been well-evaluated, it earns the chance to prove itself alongside real security work. At this stage, we actively execute solutions in parallel with working analysts in our internal SOC and Wayfinder Managed Services teams. When runs are flagged for review, an expert re-investigates the alert from scratch and establishes ground truth before reading the agent's verdict, so the reference standard can never be anchored by the machine it's meant to judge.
Every interesting divergence becomes a new case, and every fix then has to beat the case that motivated it while not regressing against the larger suite. Only after the internal teams have thoroughly vetted an agentic offering in production does it reach anyone outside SentinelOne.
Underneath all of this effort sits one standard: in security response, unmeasurable capability is unshippable capability. An agent that cannot demonstrate exactly where it is weak cannot be trusted where it is strong. Autonomy in a SOC must be earned competency by competency, evaluation by evaluation, and finally in production, in front of the people who do the job.
References
[1] G. Bernadett-Shapiro and E. Garcia Lazo. LLMs in the SOC (Part 1) | Why Benchmarks Fail Security Operations Teams. SentinelLABS, January 2026.
[2] Anthropic. Demystifying Evals for AI Agents. Anthropic Engineering, January 2026.
[3] OpenAI. Evaluation Best Practices. OpenAI API Documentation.
[4] Pydantic. Pydantic Evals. Pydantic AI Documentation.
[5] OpenAI. Advancing the Price-Performance Frontier with GPT-5.6. OpenAI, July 2026.
[6] Confident AI. LLM Agent Evaluation Metrics: Tool Calling, Task Completion, Reasoning, and Trace-Based Evals. Confident AI, 2026.
[7] R. Aleithan et al. SWE-Bench+: Enhanced Coding Benchmark for LLMs. arXiv:2410.06992, 2024.
[8] Anthropic. Writing Effective Tools for AI Agents — Using AI Agents. Anthropic Engineering, 2025.
[9] Anthropic. Building Effective AI Agents. Anthropic Research, December 2024.
[10] H. Husain. Your AI Product Needs Evals. hamel.dev, 2024.
[11] A. Szymanski et al. Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks. Proceedings of the 30th International Conference on Intelligent User Interfaces (IUI), 2025.
This article is shared at no charge for educational and informational purposes only.
Red Sky Alliance is a Cyber Threat Analysis and Intelligence Service organization. We provide indicators of compromise information (CTI) via a notification service (RedXray) or an analysis service (CTAC). For questions, comments, or assistance, please contact the office directly at 1-844-492-7225 or feedback@redskyalliance.com
- Reporting: https://www.redskyalliance.org/
- Website: https://www.redskyalliance.com/
- LinkedIn: https://www.linkedin.com/company/64265941
Weekly Cyber Intelligence Briefings:
REDSHORTS - Weekly Cyber Intelligence Briefings
https://register.gotowebinar.com/register/5207428251321676122
[1] https://www.sentinelone.com/blog/building-agents-backwards-from-evaluation/
Comments