We built two tools for evaluating AI agents on real security incidents. An agent is an AI model with tools that let it work on its own. The first tool, a sandbox, gives an agent a closed workspace with the breach logs (records of sign-ins, mail, inbox rules and files) and tools to analyze those logs. No answer key is provided. The second tool is a question generator. It creates questions and answers from the breach logs and from what the responders (the security personnel who investigated the breach) confirmed. We check the questions by hand.
The primary motivation behind this work is that, in order to secure the model weights, frontier labs need to invest in automated monitoring, reporting and response. If agents fail to correctly investigate an email breach now, this doesn't bode well for the more sophisticated attacks we're relying on them to defend against in the future.
We ran five frontier models from Anthropic and OpenAI in the sandbox on two actual business email compromise (BEC) cases, Case A and Case B. In a BEC attack, an adversary takes over an employee's email account, often to defraud the organization, for example by redirecting business payments to an account the adversary controls. For this analysis, we say an email account has been compromised when an adversary has taken control of it. A judge (a second AI model) reviewed each model's report and scored it against the findings of the responders. We held the models to the same standard as a responder. Specifically, this standard required a complete list of compromised accounts, and no innocent accounts. Leaving a compromised account off the list gives the adversary a way in. Including an innocent account sends an investigator after one of their colleagues for no reason. A report passed only if it listed every compromised account and no innocent account. One of the 30 runs passed, on Case A. On Case B, not a single run passed. The models lost most of their score on the judgment items, such as determining who was compromised. Three of the models (Claude Opus 5, Claude Sonnet 5 and GPT-5.5) lost more than a third of the possible score from Case A to Case B, while the other two (GPT-5.6-sol and gpt-6-astra) barely moved. All of this rests on two cases, five models and one attack type.
Background
Protecting trained models (their model weights, the files that make up a trained model) will take more than ensuring only approved personnel have access. Defenders will need to detect account breaches and identify when and how an attacker gained access, which accounts the attacker compromised, and which resources the attacker accessed and modified.
Determining the answers to these questions is slow and largely manual. A frontier lab is creating more telemetry than available resources can process. Telemetry contains records of sign-ins, email activity, inbox rules and file activity. So many of the designs implemented by security teams rely on automated systems: agents that reach their own conclusions without a human checking the result.
Published literature mostly demonstrates the use of language models by bad actors. Published literature demonstrating the use of AI agents to fully investigate real security incidents on their own is largely absent.
In an effort to test the capacity of AI to reach conclusions in a security investigation, we focused on the business email compromise (BEC) threat. BEC is a security incident in which an attacker takes over a company employee's email account, ultimately to commit fraud or steal data. Because we have access to real BEC case data, we are able to evaluate AI-based agents against actual cases. The evidence in a BEC case is two kinds of log: who signed in, and what happened inside each mailbox. The human security teams that responded to and investigated both incidents documented their findings, and those findings are our answer key.
Task
We give agents the same task as responders after a security incident. Responders read through the logs and document which accounts were compromised, when and how the attackers gained access to them, and what actions the attackers took. This is part of digital forensics and incident response (DFIR).
Case A and Case B are examples of business email compromise (BEC). Human responders worked these cases at the time.
The agents analyzed the logs from both cases with the same goal.
Prior to any model processing the logs, personally identifiable information (PII) was protected by obfuscation. PII includes names, email addresses, IP (Internet Protocol) addresses, organization names, and other information that, when present in logs, could identify a person or an organization. The logs provided to the models do not identify any person or organization.
There are five types of logs: sign-in logs, mailbox logs, message logs, inbox rule logs and file logs. Sign-in logs contain information about each attempt to sign in to an account. Mailbox logs capture what users did inside a mailbox. Message logs record which mail was sent and received, and between whom. Inbox rule logs capture the settings that move, hide or forward mail on their own. File logs capture which files were opened, edited or shared. Only a few accounts in the logs have been compromised.
The logs show all interactions with the mailbox, including those of an attacker. If an attacker signs in to a user's account, it is logged the same way as the user's own sign-in. Thus, the logs do not differentiate whether a sign-in was done by a legitimate user or an attacker.
The responders analyzed the logs and validated what is known as the answer key. The answer key consists of four elements: 1) the accounts which were compromised, 2) the timestamp of the attacker's initial access, 3) how the attacker accessed the account, and 4) the attacker's account activity.
Figure 1 shows the logs the agent works from and the answer key. The agent has no access to the answer key.
Figure 1. One case and its answer key. The case example on the left is a log of an actual incident. The company's systems kept five kinds of records of the incident. These records, shown in the case example, include normal traffic and records of a few accounts that were compromised and controlled by the attacker. The arrow shows the work of the responders, which led to the identification of the four facts shown on the far right. Agents do not see the responders' findings.
Related work
Benchmarks are predefined sets of questions whose answers are known, used to evaluate models and systems. A few benchmarks have been developed to evaluate language models on security tasks. For this post, we sorted them by two factors. First, is the described event something that actually happened? Second, is the task to investigate a mailbox takeover? CTIBench, AthenaBench and CyberSOCEval are based on real-world material, but focus on threat intelligence and malware analysis. No challenge, to the best of our knowledge, asked the model to analyze what happened to a mailbox
The incidents in SecRespond, SIR-Bench, DiagChain, and the Cyber Defense Benchmark are constructed in a lab or test environment, and the capture-the-flag component of DFIR-Metric uses generated challenges. The most similar work, ExCyTIn-Bench, contains two constructed business email compromises out of eight incidents. DiagChain builds its attack chains on ExCyTIn-Bench's incidents, including those two. None of these benchmarks evaluates a BEC incident using the breach logs of a real organization.
Figure 2 shows the placement of the benchmarks based on two factors: is the incident real, and is it a case of BEC? Our benchmark is the only one to evaluate a BEC case using a real-world incident.
Figure 2. Security benchmarks sorted by whether their incidents are real or staged in a lab or test tenant, and whether the task is business email compromise (BEC) or something else. The dashed red box, real incident and BEC, holds only our evaluation. ExCyTIn-Bench appears twice because it includes two staged BEC incidents and the other six are not BEC. DiagChain also appears twice, for the same reason.
Existing generators on our data
The first thing we did was test the question writing methods used in benchmarks by running them against our logs. We identified three main failure modes for this approach.
First, the answers to the questions may not be log-supported. For example, some answer keys named people who are not in our logs. Therefore, an agent that religiously reads all the logs would be marked wrong.
We could attempt to fix this by rephrasing the question. However, rephrasing does not fix an answer key that contradicts the logs.
Second, in some cases, a model may be able to infer answers to questions without examining the evidence. To test this, we generated 80 questions following a published method, hid the evidence, and had the model answer the questions. The questions are multiple choice with 5 possible answers, so there is a 20% chance of guessing right. The model matched the answer key on 70 out of 80 (87.5%), and on a second attempt on 72 out of 80 (90%). The key had answer choice 'A' as the answer 52 out of 80 times, so always selecting 'A' would match it 65% of the time. After rewriting the questions to remove these giveaways, the model matched the key only 40% of the time.
Finally, there are questions that the model cannot answer based on the logs. Even when we gave the model the answer key, it was still unable to answer them, which artificially lowers the score.
We also ran our data through the ExCyTIn-Bench generator. First, we reproduced the published result: o3 scored 0.456 on the paper's test split, as in the paper. On our logs, it generated 274 questions, of which 36 had a single verifiable answer, 116 were ambiguous, and 22 had incorrect answer keys. The method reproduces on its own data and not on ours.
Method
To create the reports, each model was run as an investigator in a closed sandbox. We created a method to generate questions and piloted it on the same cases. The method created 133 questions. We have no automatic method to confirm the accuracy of the questions generated. We, therefore, reviewed the 133 questions by hand. Questions generated through our method were not used in scoring the agents. The scores posted here come from a judge, an AI model that compares each report with the answer key. It is still undetermined whether the method will work on other cases.
Closed sandbox environment
Each model is evaluated in a closed sandbox environment, which provides it with raw telemetry and a high-level goal. A run is one attempt by one model on one case.
The model is given no worked examples, no hints as to which accounts or logs to review, and no direction while the run is in progress. The instruction provides an overview of the kinds of evidence available and the formats in which they are available. It also sets the labels the model must give each account in its report: whether the account is compromised, and if so, whether the compromise is current or historic.
The model is allowed to run Bash commands in each of its runs and interact with logs in one of three ways: Kusto Query Language (KQL), Structured Query Language (SQL), or Comma Separated Values (CSV) format. Every result comes from a real query to the full, unfiltered logs. A report is judged based on the answer key provided outside of the sandbox environment.
Figure 3 shows the complete process from the task prompt to the judged report.
Figure 3. One run: a single agent working on a single case. Once in the sandbox, the agent selects an action, and the data access layer performs that action on the unfiltered logs using Bash, SQL, KQL or CSV. The outcome is relayed back to the agent, and this process repeats for up to 150 actions or 12 hours. The final report is given to the judge, which compares it to the answer key and scores it. The answer key is made available only to the judge.
Each run is limited to 150 turns and 12 hours. One turn represents an action of the agent and the following response from the environment. The longest run required 122 turns. As turns were not documented for GPT-5.5 runs, we cannot assess if they adhered to the turn limit. The agent acts and then writes a report. The judge compares the report to the answer key
The judge is claude-opus-4-8 for the OpenAI models as well as all other models. The reports are assessed based on a rubric containing nine elements. For each element, the judge scores it three times and we take the middle score. The report's score is the weighted average of the nine element scores. The rubric is designed by us. The elements of the rubric are drawn from publicly available incident-response guidance, and describe what should be reported when a mailbox is compromised. The elements include who was compromised, the means by which the bad actor got access, what the bad actor did, the timing of the compromise, and routine facts about each mailbox. Similar elements are described in Microsoft's public guidance on compromised email accounts. A list of all elements of the rubric can be found in Figure 6. We determined the weights. Of the total weight, the element 'Who was compromised' has the highest weight, 0.30.
Checking the judge
Our judge, claude-opus-4-8, is an Anthropic model. Two of the five models it rated (i.e. Claude Opus 5 and Claude Sonnet 5) are also models of this kind. To check whether the judge showed bias with respect to its family, we carried out two checks.
We started by looking at the item of the rubric that carries the most weight, who was compromised, at 30% of the total score. Then, we compared the scores given by the judge to the F1 scores (a metric that combines precision and recall) calculated by the program for each report with respect to the account list. In general, the scores were similar. However, there were slight differences. The judge scored lower than the F1 scores for reports authored by Claude, and higher than the F1 scores for reports authored by OpenAI. To show favoritism to its family, the judge would have scored higher than the F1 for reports authored by the family model and lower than the F1 for other reports.
Second, two additional judges from OpenAI, gpt-6-sol and gpt-5.5-pro, not in the models evaluated, also rated the 30 reports using the same rubric. For these judges, we took the average change for the OpenAI reports and subtracted it from the average change for the Claude reports. In this case, a Claude judge favoring Claude would result in a negative number. We set the smallest difference we would call favoritism to 0.05. For gpt-6-sol, the difference was +0.023 (95% confidence interval -0.007 to 0.052). For gpt-5.5-pro, the difference was +0.026 (95% confidence interval -0.005 to 0.058). Both intervals include zero, and the upper ends sit just above 0.05. Both differences are positive, so compared with the OpenAI judges, the Claude judge was slightly harder on the reports from its own family.
The question generator
After human responders analyze an incident, we know what happened in it, and their confirmed findings form the answer key. Using the logs and the answer key, the question generator creates a series of questions for which it knows the answers. These questions can be answered only by going through the log data, and serve as a means to validate the models against a given set of known answers without the need for a judge to read through a report. Each answer can also be reached through a series of queries, which are stored with the question. The questions are to be answered using the given data only; they are open questions, not multiple choice. For example, a valid question can be "Did the intruder destroy anything in the mailboxes they controlled?"
Figure 4 shows a pipeline where case logs and answer keys are combined to generate questions. The first step shows the case logs in their original format, along with the facts that they can settle. The next step shows each question along with its answer and the queries that lead to it. The bottom row shows the three checks every question is put through. The first check is carried out by the system, which replays the stored queries, and is passed by 84% of the questions. The second check asks whether a question gives away its answer or fits more than one account; about 7% of the questions failed it. The last check ensures that the answer key is never provided to the system (sandbox); no leaks were found. Every question ships with any failed check recorded, and the dashed line is the path a failed question is meant to take to be rewritten.
Our question generator is able to extract the appropriate number of questions from the material without needing to artificially inflate the number of questions to meet a target. Questions that do not meet the expectations of a check are meant to be reformulated and reevaluated. Overall, the two case generated 133 questions.
Each question in the case was run through the three checks outlined in Figure 4, and the bank records the result of each check; questions that failed still ship, with the failure recorded. In check 1, the harness replays the question's stored queries and confirms whether they return the expected answer. 112 questions passed check 1. The remaining 21 questions had 3 types of issues: schema bugs, harness gaps, and a lack of data due to how the queries were written. Check 2 evaluates whether a question gives away its answer or is too general and therefore ambiguous, leading to its applicability to multiple accounts. Among the 133 questions, 9 failed check 2. Check 3 checks that the answer key is not in the sandbox, which the OS enforces. We ran controls before and after each run, and found no issues. The reason the check is enforced in the OS is that an older configuration would have leaked the answer key into the sandbox.
Results
We evaluate the runs in 2 ways. The first is a 0-to-1 score assigned by the judge, reported here. The second is a pass test on the list of accounts each model named; results from this method are included in the Evaluation criteria. Each model's score on a case is the average of its three runs, one through each interface. In general, 3 of the 5 models score much lower on Case B than on Case A: Claude Opus 5 (from 0.82 to 0.43), Claude Sonnet 5 (from 0.60 to 0.22), and GPT-5.5 (from 0.60 to 0.15). The other 2 models held their scores: GPT-5.6-sol (from 0.72 to 0.69) and gpt-6-astra (from 0.77 to 0.71).
We replicated these results when 2 OpenAI judges, identified in Checking the judge, regraded all 30 reports. The same 3 models showed the large decline, and the 2 models that held in the original evaluation also held in this follow-on evaluation. Under every judge, the decline for each of the 3 models was greater than 0.3.
Figure 5. Mean scores for 5 models for Case A (red) and Case B (blue) on the 0 to 1 scale. Bars show the mean for that model of its 3 runs on the case (using CSV, KQL and SQL). Models are in order of Case A score.
We do not know why these two models maintained their scores. One possible explanation is that they did a lot more searching than the other models. However, this is the less likely case: gpt-6-astra and GPT-5.6-sol used 21 to 37 and 40 to 122 turns respectively, so they differ widely in how much they searched.
The second possibility is that the two models were more willing to leave a question open than to fill it with a plausible name.
We can't settle this from the transcripts (exactly what each agent did and wrote), partly because of our own limitations. For the two Claude models, we recorded the summarized thinking the provider returns. For the three OpenAI models, we recorded the provider's shorter reasoning summaries, about an eighth of the text we have for the Claude runs, and none of those summaries names a compromised account. The two models that held steady are both OpenAI models, so for the models we are most interested in, we have the least available evidence.
Analysis
Where the Score was Lost
A perfect report scores a 1. On average, 0.428 of that score was lost across all runs. Figure 6 shows where the lost score is in the nine items the judge grades.
The most lost score was in three items where the model has to make a judgment instead of simply reporting, and thus, we refer to them as judgment items. In those items, the models lost score in determining who was compromised, what the attacker did, and how the attacker gained access to the system. The models assessed routine, mailbox facts the best. They got event counts and timestamps wrong fairly often; however, that particular item is weighted very low, and thus, it did not adversely affect the total score much.
Figure 6. Score lost on each of the nine rubric items judged on a report, averaged over all 30 runs (five models, two cases, three interfaces). Each element is weighted differently. The bars show the weight of the element multiplied by the average loss, so together they add up to the total loss. Items relating to judgment are shown in dark red; gray items are the remaining elements. Gray notes show common issues; the note for routine facts about mailboxes shows what went right.
The models did not fail to retrieve the data. For all models and cases, the names of the compromised accounts appeared in the models' own query results in all three runs. However, just because a name appeared in a query result does not mean the model followed up on it.
Case B has a victim whose only trace in the logs is an inbox rule. The two models that named this victim account are the two that are stable (GPT-5.6-sol and gpt-6-astra), and they named it in all three of their runs. GPT-5.5 and both the Claude models (Claude Sonnet 5 and Claude Opus 5) named the victim account in zero of their runs.
When attackers control a mailbox, they often set up an inbox rule that hides mail or directs copies of it to another mailbox. This rule persists even after the attacker is removed. Our instructions therefore told the agents that some compromises show up only as an inbox rule. Responders understand this. However, employees may have normal rules and the same rule may appear in several mailboxes, so naming a mailbox for a rule that it has may also mean naming other mailboxes as well. Most of the incorrect names were due to this: the model could not understand some rules and therefore named the mailbox as compromised. In one run, GPT-5.5 noted that several mailboxes had the same rule and still named them all as compromised.
Three models missed the Case B account, which has an inbox rule as its only trace. So the models trusted rules on accounts that were not compromised and cleared the one rule that mattered. It is likely the models are overgeneralizing a true fact. Because we never ran the cases without that line in our instructions, we are unable to assess its significance (refer to Limitations).
In all three of Claude Opus 5's Case B runs, the inbox rule appeared in its query results, and twice it noted in its reasoning that the rule looks like something an adversary might have left behind. It then searched for evidence of sign-ins and did not name the account as compromised. We have not tested this, but it seems the model needs to find evidence of attacker activity before it trusts a trace left behind by the adversary.
The recorded reasoning for GPT-5.5 does not contain anything related to this account, so we are unable to speak to that.
GPT-5.6-sol and gpt-6-astra missed compromised accounts (the victims) in Case B as well. Both models consistently missed the same account in all of their Case B runs, and in addition to that, GPT-5.6-sol also missed another account in its KQL run. No innocent account was named in any of their Case B runs.
The best run in the study was by Claude Opus 5, where it named all the victims in Case A using SQL and did not name any other account. In Case B, it missed at least two victims in every run and named an innocent account in one of its runs. GPT-5.6-sol and gpt-6-astra named no innocent accounts in any of their Case B runs. The victim the two models missed was not named in any run in the study.
Missed victims can end up in a report in one of two ways. They can be left out, or they can be listed as not compromised. The second matters more to a responder: responders tend to accept a cleared name as valid and will not check into it further. This commonly occurs. Claude Opus 5 and gpt-6-astra each listed a missed victim as not compromised in every Case B run. Likewise, Claude Sonnet 5 listed a missed victim as not compromised in two Case B runs, with GPT-5.5 and GPT-5.6-sol each listing it in a single Case B run.
False Positives
A false positive is a normal account that the agent wrongly states as compromised. In Figure 7, for each model, the figure represents the percentage of accounts the model stated that were false positives. The gray values represent the raw values for the model over its six runs. gpt-6-astra stated 17 names, with 1 being a false positive. GPT-5.5 stated 26 names, and of those, 18 were false positives.
Figure 7 shows how many of the accounts named by each model were wrong. The figure shows the share of named accounts, out of the total accounts named by the model, that were wrong. These shares are pooled over each model's six runs, three on each case.
Every wrongly named account was a real account in the organization's directory; the models invented no names. This variant of the error is more dangerous. Responders easily identify fake names and ignore them. However, real names with incorrect judgments lead responders to the wrong person.
The errors in naming accounts in the five models took two primary forms. The large majority took an inbox-rule only form. A few relied on sign-ins the models judged suspicious. None of the models named an account because the user received the phishing email, based on the volume of emails, or due to a user's job title or seniority.
Different models named the same wrong accounts. Several wrong accounts were named by more than one model, and one wrong account was named by three models. gpt-6-astra's one wrong account was also named by GPT-5.6-sol, in two of its runs. Claude Opus 5's one wrong account was also named by Claude Sonnet 5 and GPT-5.5. In both cases, the second model named the same wrong account, so it would not have caught the first model's error.
Precision and recall
Figure 7 counts only the wrong names. Figure 8 adds the missed victims. Precision and recall are defined in the Figure 8 caption. A model can be weak in one and strong in the other. In the presence of weak performance in one, F1 gives an overall performance assessment.
Figure 8 gives performance scores of the models in the form of precision, recall and F1. Precision is determined by the values shown in Figure 7. Both recall and precision can vary considerably. For example, the precision for Claude Opus 5 was 0.92, while the recall was 0.57.
Figure 8. Precision (in red), recall (in blue) and F1 (in black) of each model's list of compromised accounts, pooled over all of its runs on both cases. Values are in the range of 0 to 1. Values are calculated by code from the account lists, not by the judge. Precision is the percentage of named accounts that are actually compromised. Recall is the percentage of actually compromised accounts that appear in the model's list. F1 is the harmonic mean of precision and recall.
We consider the omission of a compromised account more harmful. If the account is not listed, the attacker keeps that account after the case is considered closed. If an account is listed but was not compromised, the analyst spends time looking into an unnecessary case. This ranking is our judgment; nothing in the data tests it.
Generally, increasing the number of named accounts increases recall and decreases precision. The five models studied, however, behaved differently from expectation. The best recall (0.76) and precision (0.94) were reported by gpt-6-astra. GPT-5.5 had the lowest precision (0.31) and the lowest recall (0.38). The other four models' recall was at or below 0.67.
The judge's score rewards finding the victims more than it penalizes naming the wrong account. In our 30 test runs, score correlated to recall at 0.91, and to precision at 0.43. In Case A, GPT-5.5 named three wrong accounts in each of two runs, and scored 0.603 and 0.654 respectively. In a different run, it named eight wrong accounts, and scored 0.550. Thus, we report precision and recall separately, because a model that names accounts freely can outscore a more careful model on the judge's single score.
Evaluation criteria
Because we don't have a human responder run to score against the same judge, we can't set a passing score for the judge's 0-1 scoring system. As such, we evaluate the runs against what is expected of a human responder: to name all the accounts that were compromised and not to name any accounts that were not compromised. Therefore, we evaluated the runs to see if they achieved 100% recall and 100% precision. Achieving both meant that a run named all the accounts that were compromised and did not name any accounts that were not compromised. A program checked both conditions for each run against the responders' findings.
Three of the 30 test runs named all the confirmed victims, and 18 of the 30 runs named no innocent accounts. One run named all the confirmed victims and no accounts that were not compromised. This run was Claude Opus 5 on Case A, through SQL. In Case B, no run by any of the models named all the confirmed victims. Twelve runs named an innocent account, and every model named an innocent account at least once, including the two models that held their score. Naming every confirmed victim is where the models fail most.
As human responders had already named all the confirmed victims, meeting this requirement would only mean that an agent could repeat a human responder's work (i.e. repeat a solved case). This doesn't mean that the agent could defend against an adversary who attacks the model weights.
Limitations
We evaluated two business email compromise cases. Our thirty runs cover those cases only. The models may need different standards to judge correct answers for other kinds of attacks (e.g. ransomware attacks) since they may leave behind different kinds of logs. We also evaluated only cases where agents acted independently. We did not evaluate cases where agents worked alongside human responders.
Each model was evaluated with each case once per interface, with no repeats. Each bar in Figure 5 shows the mean of three runs (each done using a different interface) and the variation within a bar may be relatively large (e.g. for Case A and Claude Opus 5, the three runs ranged from 0.741 to 0.972). The judge itself behaves steadily (it presents a band of approximately plus or minus 0.01 or 0.02 around the mean). Since we did not run any hypothesis testing, differences between models smaller than the variation within a bar need not be interpreted as a ranking.
The scores are generated by a language model evaluating reports against the answer key and we haven't validated that judge against a human marker. Because we haven't scored a human or a method that requires no investigation, the 0-to-1 scale has no fixed anchor. The account lists and pass standards are generated programatically. The programs do not rely on the judge. We checked whether the judge favored its own model family and found no sign that it did (see Checking the Judge, in the Method section). No human evaluator has checked the three judges, and so it's possible the judges all share the same preference. We have no means to share our logs and so there is no means for the public to analyze or validate our results.
Agents were instructed to tell the difference between a legitimate user's sign in and an attacker's by considering various attributes including the location, device and time. Additionally, agents were instructed to remember certain attacks leave no trace of the compromise except an inbox rule. We did not evaluate models in the absence of the above heuristics and so it is not possible to state what effect, if any, the heuristics had on the models. We can state when an answer was incorrect, however, the reason the model got it wrong is often unclear (see Results).
Discussion
The important part is not how many points a model loses. The more interesting point is where a model loses points. The models lost points interpreting the evidence. The models retrieve evidence and bring it into their context window, but they have issues understanding the logic behind it. Models get points for retrieving evidence but lose points on understanding the bigger picture of which the evidence is a part. Improvements to the context window or the speed of log search would help the models retrieve and present the evidence, which they are already able to do. Models would not benefit a great deal from this on the parts they struggle with.
As mentioned in the Analysis, all compromised accounts appeared in the results of the models' queries. Therefore, the models were not performing blind searches. We read the recorded reasoning of Claude's six runs on Case B. In five of the six, the model considered the account whose only evidence was an inbox rule, and wrote a reason to clear it. In the remaining run (Claude Sonnet 5, using KQL), the model explained it would need to perform a lookup to retrieve evidence, and never did. This provides some support to our guess that Claude Opus 5 needed evidence of an attacker to trust the rule. The recorded reasoning is a summary, and beyond this, we cannot conclude anything.
In both cases, the models had the records in front of them, and most of the points they lost were lost in interpreting the records.
We are least certain about applying these observations to the defense of model weights. One line of reasoning is as follows. First, in comparison to a BEC attack, we expect that model weight theft would require higher levels of sophistication in terms of covertness. Second, model weight theft-related logs would likely be more disordered. Third, no responders would have figured out the answer to the case before the model does. Consequently, the failure modes seen here would be even more concerning. Of course, it would be too speculative to make such statements based on these two cases. In our data set, 2 of the 5 models demonstrated very little score loss between Case A and Case B. If the models were catching up with the task, then we would expect this. Thus, for the purpose of this post, we will not be taking this line of reasoning into consideration.
The results help identify the most effective use of resources. A security plan that involves agents investigating security breaches would want to test it on a case that has already been solved by responders, as the answer key for that case has already been confirmed. Three questions remain open. The first question is, in the context of verification, can we rely on it enough to completely remove the human in the loop? The second question is, how much would the generator improve if each of its mistakes were caught and used to train it? The last question is, would the results hold for other cases and other models?
These two cases do not truly answer the question "Can agents be trusted to investigate cases of a security breach?" However, they show that the five models examined fail on two closed cases, and indicate where the models fail. True evaluation of this trust will require more cases, different types of attacks and a human baseline, all scored by the same criteria. The models will eventually solve these two cases. The generator has only been tested on two incidents, and each of its questions has been reviewed by hand. The main constraint is the lack of cases. If you are working on similar problems, have cases available, and/or are interested in working on this collaboration, please contact us using the information below.
Acknowledgments
I would like to express my sincere gratitude to my mentors Alexis Carlier, Zainab Ali Majid, Alex Chan, Jeewoo Kim and Thomas Morris for their support and guidance during my research, which was performed under the MATS program (Summer 2026 cohort). For questions or collaboration, email hemantkumarbk@arizona.edu
TL;DR
We built two tools for evaluating AI agents on real security incidents. An agent is an AI model with tools that let it work on its own. The first tool, a sandbox, gives an agent a closed workspace with the breach logs (records of sign-ins, mail, inbox rules and files) and tools to analyze those logs. No answer key is provided. The second tool is a question generator. It creates questions and answers from the breach logs and from what the responders (the security personnel who investigated the breach) confirmed. We check the questions by hand.
The primary motivation behind this work is that, in order to secure the model weights, frontier labs need to invest in automated monitoring, reporting and response. If agents fail to correctly investigate an email breach now, this doesn't bode well for the more sophisticated attacks we're relying on them to defend against in the future.
We ran five frontier models from Anthropic and OpenAI in the sandbox on two actual business email compromise (BEC) cases, Case A and Case B. In a BEC attack, an adversary takes over an employee's email account, often to defraud the organization, for example by redirecting business payments to an account the adversary controls. For this analysis, we say an email account has been compromised when an adversary has taken control of it. A judge (a second AI model) reviewed each model's report and scored it against the findings of the responders. We held the models to the same standard as a responder. Specifically, this standard required a complete list of compromised accounts, and no innocent accounts. Leaving a compromised account off the list gives the adversary a way in. Including an innocent account sends an investigator after one of their colleagues for no reason. A report passed only if it listed every compromised account and no innocent account. One of the 30 runs passed, on Case A. On Case B, not a single run passed. The models lost most of their score on the judgment items, such as determining who was compromised. Three of the models (Claude Opus 5, Claude Sonnet 5 and GPT-5.5) lost more than a third of the possible score from Case A to Case B, while the other two (GPT-5.6-sol and gpt-6-astra) barely moved. All of this rests on two cases, five models and one attack type.
Background
Protecting trained models (their model weights, the files that make up a trained model) will take more than ensuring only approved personnel have access. Defenders will need to detect account breaches and identify when and how an attacker gained access, which accounts the attacker compromised, and which resources the attacker accessed and modified.
Determining the answers to these questions is slow and largely manual. A frontier lab is creating more telemetry than available resources can process. Telemetry contains records of sign-ins, email activity, inbox rules and file activity. So many of the designs implemented by security teams rely on automated systems: agents that reach their own conclusions without a human checking the result.
Published literature mostly demonstrates the use of language models by bad actors. Published literature demonstrating the use of AI agents to fully investigate real security incidents on their own is largely absent.
In an effort to test the capacity of AI to reach conclusions in a security investigation, we focused on the business email compromise (BEC) threat. BEC is a security incident in which an attacker takes over a company employee's email account, ultimately to commit fraud or steal data. Because we have access to real BEC case data, we are able to evaluate AI-based agents against actual cases. The evidence in a BEC case is two kinds of log: who signed in, and what happened inside each mailbox. The human security teams that responded to and investigated both incidents documented their findings, and those findings are our answer key.
Task
We give agents the same task as responders after a security incident. Responders read through the logs and document which accounts were compromised, when and how the attackers gained access to them, and what actions the attackers took. This is part of digital forensics and incident response (DFIR).
Case A and Case B are examples of business email compromise (BEC). Human responders worked these cases at the time.
The agents analyzed the logs from both cases with the same goal.
Prior to any model processing the logs, personally identifiable information (PII) was protected by obfuscation. PII includes names, email addresses, IP (Internet Protocol) addresses, organization names, and other information that, when present in logs, could identify a person or an organization. The logs provided to the models do not identify any person or organization.
There are five types of logs: sign-in logs, mailbox logs, message logs, inbox rule logs and file logs. Sign-in logs contain information about each attempt to sign in to an account. Mailbox logs capture what users did inside a mailbox. Message logs record which mail was sent and received, and between whom. Inbox rule logs capture the settings that move, hide or forward mail on their own. File logs capture which files were opened, edited or shared. Only a few accounts in the logs have been compromised.
The logs show all interactions with the mailbox, including those of an attacker. If an attacker signs in to a user's account, it is logged the same way as the user's own sign-in. Thus, the logs do not differentiate whether a sign-in was done by a legitimate user or an attacker.
The responders analyzed the logs and validated what is known as the answer key. The answer key consists of four elements: 1) the accounts which were compromised, 2) the timestamp of the attacker's initial access, 3) how the attacker accessed the account, and 4) the attacker's account activity.
Figure 1 shows the logs the agent works from and the answer key. The agent has no access to the answer key.
Figure 1. One case and its answer key. The case example on the left is a log of an actual incident. The company's systems kept five kinds of records of the incident. These records, shown in the case example, include normal traffic and records of a few accounts that were compromised and controlled by the attacker. The arrow shows the work of the responders, which led to the identification of the four facts shown on the far right. Agents do not see the responders' findings.
Related work
Benchmarks are predefined sets of questions whose answers are known, used to evaluate models and systems. A few benchmarks have been developed to evaluate language models on security tasks. For this post, we sorted them by two factors. First, is the described event something that actually happened? Second, is the task to investigate a mailbox takeover? CTIBench, AthenaBench and CyberSOCEval are based on real-world material, but focus on threat intelligence and malware analysis. No challenge, to the best of our knowledge, asked the model to analyze what happened to a mailbox
The incidents in SecRespond, SIR-Bench, DiagChain, and the Cyber Defense Benchmark are constructed in a lab or test environment, and the capture-the-flag component of DFIR-Metric uses generated challenges. The most similar work, ExCyTIn-Bench, contains two constructed business email compromises out of eight incidents. DiagChain builds its attack chains on ExCyTIn-Bench's incidents, including those two. None of these benchmarks evaluates a BEC incident using the breach logs of a real organization.
Figure 2 shows the placement of the benchmarks based on two factors: is the incident real, and is it a case of BEC? Our benchmark is the only one to evaluate a BEC case using a real-world incident.
Figure 2. Security benchmarks sorted by whether their incidents are real or staged in a lab or test tenant, and whether the task is business email compromise (BEC) or something else. The dashed red box, real incident and BEC, holds only our evaluation. ExCyTIn-Bench appears twice because it includes two staged BEC incidents and the other six are not BEC. DiagChain also appears twice, for the same reason.
Existing generators on our data
The first thing we did was test the question writing methods used in benchmarks by running them against our logs. We identified three main failure modes for this approach.
First, the answers to the questions may not be log-supported. For example, some answer keys named people who are not in our logs. Therefore, an agent that religiously reads all the logs would be marked wrong.
We could attempt to fix this by rephrasing the question. However, rephrasing does not fix an answer key that contradicts the logs.
Second, in some cases, a model may be able to infer answers to questions without examining the evidence. To test this, we generated 80 questions following a published method, hid the evidence, and had the model answer the questions. The questions are multiple choice with 5 possible answers, so there is a 20% chance of guessing right. The model matched the answer key on 70 out of 80 (87.5%), and on a second attempt on 72 out of 80 (90%). The key had answer choice 'A' as the answer 52 out of 80 times, so always selecting 'A' would match it 65% of the time. After rewriting the questions to remove these giveaways, the model matched the key only 40% of the time.
Finally, there are questions that the model cannot answer based on the logs. Even when we gave the model the answer key, it was still unable to answer them, which artificially lowers the score.
We also ran our data through the ExCyTIn-Bench generator. First, we reproduced the published result: o3 scored 0.456 on the paper's test split, as in the paper. On our logs, it generated 274 questions, of which 36 had a single verifiable answer, 116 were ambiguous, and 22 had incorrect answer keys. The method reproduces on its own data and not on ours.
Method
To create the reports, each model was run as an investigator in a closed sandbox. We created a method to generate questions and piloted it on the same cases. The method created 133 questions. We have no automatic method to confirm the accuracy of the questions generated. We, therefore, reviewed the 133 questions by hand. Questions generated through our method were not used in scoring the agents. The scores posted here come from a judge, an AI model that compares each report with the answer key. It is still undetermined whether the method will work on other cases.
Closed sandbox environment
Each model is evaluated in a closed sandbox environment, which provides it with raw telemetry and a high-level goal. A run is one attempt by one model on one case.
The model is given no worked examples, no hints as to which accounts or logs to review, and no direction while the run is in progress. The instruction provides an overview of the kinds of evidence available and the formats in which they are available. It also sets the labels the model must give each account in its report: whether the account is compromised, and if so, whether the compromise is current or historic.
The model is allowed to run Bash commands in each of its runs and interact with logs in one of three ways: Kusto Query Language (KQL), Structured Query Language (SQL), or Comma Separated Values (CSV) format. Every result comes from a real query to the full, unfiltered logs. A report is judged based on the answer key provided outside of the sandbox environment.
Figure 3 shows the complete process from the task prompt to the judged report.
Figure 3. One run: a single agent working on a single case. Once in the sandbox, the agent selects an action, and the data access layer performs that action on the unfiltered logs using Bash, SQL, KQL or CSV. The outcome is relayed back to the agent, and this process repeats for up to 150 actions or 12 hours. The final report is given to the judge, which compares it to the answer key and scores it. The answer key is made available only to the judge.
Each run is limited to 150 turns and 12 hours. One turn represents an action of the agent and the following response from the environment. The longest run required 122 turns. As turns were not documented for GPT-5.5 runs, we cannot assess if they adhered to the turn limit. The agent acts and then writes a report. The judge compares the report to the answer key
The judge is claude-opus-4-8 for the OpenAI models as well as all other models. The reports are assessed based on a rubric containing nine elements. For each element, the judge scores it three times and we take the middle score. The report's score is the weighted average of the nine element scores. The rubric is designed by us. The elements of the rubric are drawn from publicly available incident-response guidance, and describe what should be reported when a mailbox is compromised. The elements include who was compromised, the means by which the bad actor got access, what the bad actor did, the timing of the compromise, and routine facts about each mailbox. Similar elements are described in Microsoft's public guidance on compromised email accounts. A list of all elements of the rubric can be found in Figure 6. We determined the weights. Of the total weight, the element 'Who was compromised' has the highest weight, 0.30.
Checking the judge
Our judge, claude-opus-4-8, is an Anthropic model. Two of the five models it rated (i.e. Claude Opus 5 and Claude Sonnet 5) are also models of this kind. To check whether the judge showed bias with respect to its family, we carried out two checks.
We started by looking at the item of the rubric that carries the most weight, who was compromised, at 30% of the total score. Then, we compared the scores given by the judge to the F1 scores (a metric that combines precision and recall) calculated by the program for each report with respect to the account list. In general, the scores were similar. However, there were slight differences. The judge scored lower than the F1 scores for reports authored by Claude, and higher than the F1 scores for reports authored by OpenAI. To show favoritism to its family, the judge would have scored higher than the F1 for reports authored by the family model and lower than the F1 for other reports.
Second, two additional judges from OpenAI, gpt-6-sol and gpt-5.5-pro, not in the models evaluated, also rated the 30 reports using the same rubric. For these judges, we took the average change for the OpenAI reports and subtracted it from the average change for the Claude reports. In this case, a Claude judge favoring Claude would result in a negative number. We set the smallest difference we would call favoritism to 0.05. For gpt-6-sol, the difference was +0.023 (95% confidence interval -0.007 to 0.052). For gpt-5.5-pro, the difference was +0.026 (95% confidence interval -0.005 to 0.058). Both intervals include zero, and the upper ends sit just above 0.05. Both differences are positive, so compared with the OpenAI judges, the Claude judge was slightly harder on the reports from its own family.
The question generator
After human responders analyze an incident, we know what happened in it, and their confirmed findings form the answer key. Using the logs and the answer key, the question generator creates a series of questions for which it knows the answers. These questions can be answered only by going through the log data, and serve as a means to validate the models against a given set of known answers without the need for a judge to read through a report. Each answer can also be reached through a series of queries, which are stored with the question. The questions are to be answered using the given data only; they are open questions, not multiple choice. For example, a valid question can be "Did the intruder destroy anything in the mailboxes they controlled?"
Figure 4 shows a pipeline where case logs and answer keys are combined to generate questions. The first step shows the case logs in their original format, along with the facts that they can settle. The next step shows each question along with its answer and the queries that lead to it. The bottom row shows the three checks every question is put through. The first check is carried out by the system, which replays the stored queries, and is passed by 84% of the questions. The second check asks whether a question gives away its answer or fits more than one account; about 7% of the questions failed it. The last check ensures that the answer key is never provided to the system (sandbox); no leaks were found. Every question ships with any failed check recorded, and the dashed line is the path a failed question is meant to take to be rewritten.
Our question generator is able to extract the appropriate number of questions from the material without needing to artificially inflate the number of questions to meet a target. Questions that do not meet the expectations of a check are meant to be reformulated and reevaluated. Overall, the two case generated 133 questions.
Each question in the case was run through the three checks outlined in Figure 4, and the bank records the result of each check; questions that failed still ship, with the failure recorded. In check 1, the harness replays the question's stored queries and confirms whether they return the expected answer. 112 questions passed check 1. The remaining 21 questions had 3 types of issues: schema bugs, harness gaps, and a lack of data due to how the queries were written. Check 2 evaluates whether a question gives away its answer or is too general and therefore ambiguous, leading to its applicability to multiple accounts. Among the 133 questions, 9 failed check 2. Check 3 checks that the answer key is not in the sandbox, which the OS enforces. We ran controls before and after each run, and found no issues. The reason the check is enforced in the OS is that an older configuration would have leaked the answer key into the sandbox.
Results
We evaluate the runs in 2 ways. The first is a 0-to-1 score assigned by the judge, reported here. The second is a pass test on the list of accounts each model named; results from this method are included in the Evaluation criteria. Each model's score on a case is the average of its three runs, one through each interface. In general, 3 of the 5 models score much lower on Case B than on Case A: Claude Opus 5 (from 0.82 to 0.43), Claude Sonnet 5 (from 0.60 to 0.22), and GPT-5.5 (from 0.60 to 0.15). The other 2 models held their scores: GPT-5.6-sol (from 0.72 to 0.69) and gpt-6-astra (from 0.77 to 0.71).
We replicated these results when 2 OpenAI judges, identified in Checking the judge, regraded all 30 reports. The same 3 models showed the large decline, and the 2 models that held in the original evaluation also held in this follow-on evaluation. Under every judge, the decline for each of the 3 models was greater than 0.3.
Figure 5. Mean scores for 5 models for Case A (red) and Case B (blue) on the 0 to 1 scale. Bars show the mean for that model of its 3 runs on the case (using CSV, KQL and SQL). Models are in order of Case A score.
We do not know why these two models maintained their scores. One possible explanation is that they did a lot more searching than the other models. However, this is the less likely case: gpt-6-astra and GPT-5.6-sol used 21 to 37 and 40 to 122 turns respectively, so they differ widely in how much they searched.
The second possibility is that the two models were more willing to leave a question open than to fill it with a plausible name.
We can't settle this from the transcripts (exactly what each agent did and wrote), partly because of our own limitations. For the two Claude models, we recorded the summarized thinking the provider returns. For the three OpenAI models, we recorded the provider's shorter reasoning summaries, about an eighth of the text we have for the Claude runs, and none of those summaries names a compromised account. The two models that held steady are both OpenAI models, so for the models we are most interested in, we have the least available evidence.
Analysis
Where the Score was Lost
A perfect report scores a 1. On average, 0.428 of that score was lost across all runs. Figure 6 shows where the lost score is in the nine items the judge grades.
The most lost score was in three items where the model has to make a judgment instead of simply reporting, and thus, we refer to them as judgment items. In those items, the models lost score in determining who was compromised, what the attacker did, and how the attacker gained access to the system. The models assessed routine, mailbox facts the best. They got event counts and timestamps wrong fairly often; however, that particular item is weighted very low, and thus, it did not adversely affect the total score much.
Figure 6. Score lost on each of the nine rubric items judged on a report, averaged over all 30 runs (five models, two cases, three interfaces). Each element is weighted differently. The bars show the weight of the element multiplied by the average loss, so together they add up to the total loss. Items relating to judgment are shown in dark red; gray items are the remaining elements. Gray notes show common issues; the note for routine facts about mailboxes shows what went right.
The models did not fail to retrieve the data. For all models and cases, the names of the compromised accounts appeared in the models' own query results in all three runs. However, just because a name appeared in a query result does not mean the model followed up on it.
Case B has a victim whose only trace in the logs is an inbox rule. The two models that named this victim account are the two that are stable (GPT-5.6-sol and gpt-6-astra), and they named it in all three of their runs. GPT-5.5 and both the Claude models (Claude Sonnet 5 and Claude Opus 5) named the victim account in zero of their runs.
When attackers control a mailbox, they often set up an inbox rule that hides mail or directs copies of it to another mailbox. This rule persists even after the attacker is removed. Our instructions therefore told the agents that some compromises show up only as an inbox rule. Responders understand this. However, employees may have normal rules and the same rule may appear in several mailboxes, so naming a mailbox for a rule that it has may also mean naming other mailboxes as well. Most of the incorrect names were due to this: the model could not understand some rules and therefore named the mailbox as compromised. In one run, GPT-5.5 noted that several mailboxes had the same rule and still named them all as compromised.
Three models missed the Case B account, which has an inbox rule as its only trace. So the models trusted rules on accounts that were not compromised and cleared the one rule that mattered. It is likely the models are overgeneralizing a true fact. Because we never ran the cases without that line in our instructions, we are unable to assess its significance (refer to Limitations).
In all three of Claude Opus 5's Case B runs, the inbox rule appeared in its query results, and twice it noted in its reasoning that the rule looks like something an adversary might have left behind. It then searched for evidence of sign-ins and did not name the account as compromised. We have not tested this, but it seems the model needs to find evidence of attacker activity before it trusts a trace left behind by the adversary.
The recorded reasoning for GPT-5.5 does not contain anything related to this account, so we are unable to speak to that.
GPT-5.6-sol and gpt-6-astra missed compromised accounts (the victims) in Case B as well. Both models consistently missed the same account in all of their Case B runs, and in addition to that, GPT-5.6-sol also missed another account in its KQL run. No innocent account was named in any of their Case B runs.
The best run in the study was by Claude Opus 5, where it named all the victims in Case A using SQL and did not name any other account. In Case B, it missed at least two victims in every run and named an innocent account in one of its runs. GPT-5.6-sol and gpt-6-astra named no innocent accounts in any of their Case B runs. The victim the two models missed was not named in any run in the study.
Missed victims can end up in a report in one of two ways. They can be left out, or they can be listed as not compromised. The second matters more to a responder: responders tend to accept a cleared name as valid and will not check into it further. This commonly occurs. Claude Opus 5 and gpt-6-astra each listed a missed victim as not compromised in every Case B run. Likewise, Claude Sonnet 5 listed a missed victim as not compromised in two Case B runs, with GPT-5.5 and GPT-5.6-sol each listing it in a single Case B run.
False Positives
A false positive is a normal account that the agent wrongly states as compromised. In Figure 7, for each model, the figure represents the percentage of accounts the model stated that were false positives. The gray values represent the raw values for the model over its six runs. gpt-6-astra stated 17 names, with 1 being a false positive. GPT-5.5 stated 26 names, and of those, 18 were false positives.
Figure 7 shows how many of the accounts named by each model were wrong. The figure shows the share of named accounts, out of the total accounts named by the model, that were wrong. These shares are pooled over each model's six runs, three on each case.
Every wrongly named account was a real account in the organization's directory; the models invented no names. This variant of the error is more dangerous. Responders easily identify fake names and ignore them. However, real names with incorrect judgments lead responders to the wrong person.
The errors in naming accounts in the five models took two primary forms. The large majority took an inbox-rule only form. A few relied on sign-ins the models judged suspicious. None of the models named an account because the user received the phishing email, based on the volume of emails, or due to a user's job title or seniority.
Different models named the same wrong accounts. Several wrong accounts were named by more than one model, and one wrong account was named by three models. gpt-6-astra's one wrong account was also named by GPT-5.6-sol, in two of its runs. Claude Opus 5's one wrong account was also named by Claude Sonnet 5 and GPT-5.5. In both cases, the second model named the same wrong account, so it would not have caught the first model's error.
Precision and recall
Figure 7 counts only the wrong names. Figure 8 adds the missed victims. Precision and recall are defined in the Figure 8 caption. A model can be weak in one and strong in the other. In the presence of weak performance in one, F1 gives an overall performance assessment.
Figure 8 gives performance scores of the models in the form of precision, recall and F1. Precision is determined by the values shown in Figure 7. Both recall and precision can vary considerably. For example, the precision for Claude Opus 5 was 0.92, while the recall was 0.57.
Figure 8. Precision (in red), recall (in blue) and F1 (in black) of each model's list of compromised accounts, pooled over all of its runs on both cases. Values are in the range of 0 to 1. Values are calculated by code from the account lists, not by the judge. Precision is the percentage of named accounts that are actually compromised. Recall is the percentage of actually compromised accounts that appear in the model's list. F1 is the harmonic mean of precision and recall.
We consider the omission of a compromised account more harmful. If the account is not listed, the attacker keeps that account after the case is considered closed. If an account is listed but was not compromised, the analyst spends time looking into an unnecessary case. This ranking is our judgment; nothing in the data tests it.
Generally, increasing the number of named accounts increases recall and decreases precision. The five models studied, however, behaved differently from expectation. The best recall (0.76) and precision (0.94) were reported by gpt-6-astra. GPT-5.5 had the lowest precision (0.31) and the lowest recall (0.38). The other four models' recall was at or below 0.67.
The judge's score rewards finding the victims more than it penalizes naming the wrong account. In our 30 test runs, score correlated to recall at 0.91, and to precision at 0.43. In Case A, GPT-5.5 named three wrong accounts in each of two runs, and scored 0.603 and 0.654 respectively. In a different run, it named eight wrong accounts, and scored 0.550. Thus, we report precision and recall separately, because a model that names accounts freely can outscore a more careful model on the judge's single score.
Evaluation criteria
Because we don't have a human responder run to score against the same judge, we can't set a passing score for the judge's 0-1 scoring system. As such, we evaluate the runs against what is expected of a human responder: to name all the accounts that were compromised and not to name any accounts that were not compromised. Therefore, we evaluated the runs to see if they achieved 100% recall and 100% precision. Achieving both meant that a run named all the accounts that were compromised and did not name any accounts that were not compromised. A program checked both conditions for each run against the responders' findings.
Three of the 30 test runs named all the confirmed victims, and 18 of the 30 runs named no innocent accounts. One run named all the confirmed victims and no accounts that were not compromised. This run was Claude Opus 5 on Case A, through SQL. In Case B, no run by any of the models named all the confirmed victims. Twelve runs named an innocent account, and every model named an innocent account at least once, including the two models that held their score. Naming every confirmed victim is where the models fail most.
As human responders had already named all the confirmed victims, meeting this requirement would only mean that an agent could repeat a human responder's work (i.e. repeat a solved case). This doesn't mean that the agent could defend against an adversary who attacks the model weights.
Limitations
We evaluated two business email compromise cases. Our thirty runs cover those cases only. The models may need different standards to judge correct answers for other kinds of attacks (e.g. ransomware attacks) since they may leave behind different kinds of logs. We also evaluated only cases where agents acted independently. We did not evaluate cases where agents worked alongside human responders.
Each model was evaluated with each case once per interface, with no repeats. Each bar in Figure 5 shows the mean of three runs (each done using a different interface) and the variation within a bar may be relatively large (e.g. for Case A and Claude Opus 5, the three runs ranged from 0.741 to 0.972). The judge itself behaves steadily (it presents a band of approximately plus or minus 0.01 or 0.02 around the mean). Since we did not run any hypothesis testing, differences between models smaller than the variation within a bar need not be interpreted as a ranking.
The scores are generated by a language model evaluating reports against the answer key and we haven't validated that judge against a human marker. Because we haven't scored a human or a method that requires no investigation, the 0-to-1 scale has no fixed anchor. The account lists and pass standards are generated programatically. The programs do not rely on the judge. We checked whether the judge favored its own model family and found no sign that it did (see Checking the Judge, in the Method section). No human evaluator has checked the three judges, and so it's possible the judges all share the same preference. We have no means to share our logs and so there is no means for the public to analyze or validate our results.
Agents were instructed to tell the difference between a legitimate user's sign in and an attacker's by considering various attributes including the location, device and time. Additionally, agents were instructed to remember certain attacks leave no trace of the compromise except an inbox rule. We did not evaluate models in the absence of the above heuristics and so it is not possible to state what effect, if any, the heuristics had on the models. We can state when an answer was incorrect, however, the reason the model got it wrong is often unclear (see Results).
Discussion
The important part is not how many points a model loses. The more interesting point is where a model loses points. The models lost points interpreting the evidence. The models retrieve evidence and bring it into their context window, but they have issues understanding the logic behind it. Models get points for retrieving evidence but lose points on understanding the bigger picture of which the evidence is a part. Improvements to the context window or the speed of log search would help the models retrieve and present the evidence, which they are already able to do. Models would not benefit a great deal from this on the parts they struggle with.
As mentioned in the Analysis, all compromised accounts appeared in the results of the models' queries. Therefore, the models were not performing blind searches. We read the recorded reasoning of Claude's six runs on Case B. In five of the six, the model considered the account whose only evidence was an inbox rule, and wrote a reason to clear it. In the remaining run (Claude Sonnet 5, using KQL), the model explained it would need to perform a lookup to retrieve evidence, and never did. This provides some support to our guess that Claude Opus 5 needed evidence of an attacker to trust the rule. The recorded reasoning is a summary, and beyond this, we cannot conclude anything.
In both cases, the models had the records in front of them, and most of the points they lost were lost in interpreting the records.
We are least certain about applying these observations to the defense of model weights. One line of reasoning is as follows. First, in comparison to a BEC attack, we expect that model weight theft would require higher levels of sophistication in terms of covertness. Second, model weight theft-related logs would likely be more disordered. Third, no responders would have figured out the answer to the case before the model does. Consequently, the failure modes seen here would be even more concerning. Of course, it would be too speculative to make such statements based on these two cases. In our data set, 2 of the 5 models demonstrated very little score loss between Case A and Case B. If the models were catching up with the task, then we would expect this. Thus, for the purpose of this post, we will not be taking this line of reasoning into consideration.
The results help identify the most effective use of resources. A security plan that involves agents investigating security breaches would want to test it on a case that has already been solved by responders, as the answer key for that case has already been confirmed. Three questions remain open. The first question is, in the context of verification, can we rely on it enough to completely remove the human in the loop? The second question is, how much would the generator improve if each of its mistakes were caught and used to train it? The last question is, would the results hold for other cases and other models?
These two cases do not truly answer the question "Can agents be trusted to investigate cases of a security breach?" However, they show that the five models examined fail on two closed cases, and indicate where the models fail. True evaluation of this trust will require more cases, different types of attacks and a human baseline, all scored by the same criteria. The models will eventually solve these two cases. The generator has only been tested on two incidents, and each of its questions has been reviewed by hand. The main constraint is the lack of cases. If you are working on similar problems, have cases available, and/or are interested in working on this collaboration, please contact us using the information below.
Acknowledgments
I would like to express my sincere gratitude to my mentors Alexis Carlier, Zainab Ali Majid, Alex Chan, Jeewoo Kim and Thomas Morris for their support and guidance during my research, which was performed under the MATS program (Summer 2026 cohort). For questions or collaboration, email hemantkumarbk@arizona.edu