This is an automated rejection. No LLM generated, assisted/co-written, or edited work.
Read full explanation
As someone who has done a lot of machine learning projects, starting from the EDA step to deploying the model, it has started to feel quite repetitive. Different projects have their specific nuances and difficulties, but a brunt of the work is part of the same ML life-cycle, and a lot of the work I have been doing recently has been more or less on autopilot. I thought if it feels like I am on autopilot while doing this work, then these frontier agentic models would be great at automating this process. This is not some genius idea that no one has had before, especially with frameworks like AutoML being built even before mass AI integration into the data science workflow. But my focus isn't just automating this process and getting the best possible metrics; it's focused on trying to create a interpretable tool that catches hallucination, giving the user complete understanding and reasoning for every decision made.
The most pressing issue today in the field of AI is, as we give these agents more freedom in our daily lives, and as the problems they work on get more advanced, will we get left behind intelligence-wise on how they completed these tasks. Or even worse, will the measures they take to solve the tasks we give them have unintended consequences? This has been especially brought to light with the OpenAI-Hugging Face incident, in which agents running an internal OpenAI capability evaluation escaped their sandbox, moved through OpenAI's internal network, and eventually reached Hugging Face's production infrastructure, hunting for an answer key to the benchmark they were being evaluated on. Bad enough as the sandbox escape is, it's a detail from OpenAI's own report that has caught everyone's attention: one agent recognized the action was "arguably unauthorized" and paused, until another agent posted "GO" on an internal coordination channel with a deadline, and the first agent proceeded, reasoning the "GO authorization arrived." Authority transferred unknown from these OpenAI researchers who assigned the task to whatever the other agents said.
This is why it is more important to build more interpretable and aligned models than simply the most powerful model possible. This is the kind of system I tried to build with this project, and I think it is very important as traditional machine learning is being employed more and more in different fields such as healthcare, where a misrepresented statistic or a hallucination could have real impact on people's lives.
How the pipeline works
The pipeline has four stages. First, there's an EDA agent that ingests the uploaded dataset and goes through each column, making observations on important properties of each feature such as skewness and multicollinearity. Then there's the preprocessing agent, which decides whether to impute, log-transform, drop, or scale each column. Then there's the model selection agent which screens a fixed set of candidate models one at a time, accepting or rejecting each based on whether it fits the dataset's type and shape (categorical vs. regression, feature cardinality, sample size). Finally, there's the model evaluation agent, which reports the top-ranked accepted model's performance through metrics and reasoning.
┌─────────────────────────────────────────────┐ │ EDA → Preprocessing → Model Select → Eval │ └───────────────────┬───────────────────────────┘ │ every decision = a Claim │ citing a specific EvidenceRef ▼ ┌──────────────────────────┐ │ Layer 1 (deterministic) │ │ Is the cited value real │ │ and accurate? │ └────────────┬─────────────────┘ │ if accurate ▼ ┌──────────────────────────┐ │ Layer 2 (LLM critic, │ │ isolated context) │ │ Does the reasoning │ │ from that evidence │ │ actually hold up? │ │ Scored 0-3 │ └──────────────────────────┘
Each of these agents makes a multitude of individual claims during each stage explaining their reason for dropping a column or selecting a model. Before any agent runs, a separate deterministic script computes a fixed set of statistics for every column in the dataset, such as missingness or correlation with the target. These values are the backbone of the entire pipeline as every decision and its reasoning has to be grounded in one of them. This is deliberate so we can force every decision to cite a checkable number to keep the pipeline trustworthy and limit hallucination. No claim is allowed to just assert something about the data as it has to point to a real computed number. Here's what one actually looks like that was pulled directly from one of the runs (a leaked column injected into the Titanic dataset for a stress test, covered later in this post):
{ "claim_id": "eda_001", "decision": "Remove SYNTH_LEAKED: extreme target leakage detected", "evidence": [ {"metric_key": "col:SYNTH_LEAKED:leakage_suspect", "claimed_value": true}, {"metric_key": "col:SYNTH_LEAKED:corr_with_target", "claimed_value": 0.99} ], "evidence_verdicts": [ {"metric_key": "col:SYNTH_LEAKED:leakage_suspect", "layer1_code": "EVIDENCE_ACCURATE"}, {"metric_key": "col:SYNTH_LEAKED:corr_with_target", "layer1_code": "EVIDENCE_ACCURATE"} ], "justification": "SYNTH_LEAKED shows a correlation of 0.99 with the target, indicating near-perfect predictive power that would not generalize. This is a clear data leakage issue.", "severity": "critical", "layer2_score": 3 }
The evidence array is a specific metric key and value the agent cited. The evidence_verdicts array is the Layer 1 critic's deterministic check that those values exist. The layer2_score is the Layer 2 critic's independent judgment, in an isolated context that never sees the rest of the agent's conversation, of whether the reasoning connecting that evidence to the decision makes sense.
The first layer checks each claim's citation deterministically: does the claimed value actually match what was computed, or did the agent fabricate a number, or cite a real metric but misstate its value. These are tracked as two separate failure types.
The second layer is an LLM critic that never sees the full agent conversation, only the individual claim and its cited evidence. That critic judges how logical the reasoning connecting the evidence to the decision is by scoring it on a 0-3 scale. A claim can cite a perfectly real number and still reason badly and this type of mistake is supposed to be caught by this layer.
To a paint a clear picture of what to expect when using the pipeline, the input is just the raw dataset with the selected target column and a task type (classification, regression, etc.) but with no dataset description or domain context needed. The output is the full decision record: every claim made at every stage, the evidence it cited, whether that citation checked out, and whether the reasoning behind it held up along with the trained model and its metrics.
Dataset Overview
I chose to test and evaluate the pipeline using five separate datasets, each chosen to exercise a different part of the pipeline: Titanic (heavy missingness, stress-tests imputation reasoning), Wine Quality (multiclass), Adult Income (binary, roughly balanced), Ames Housing (regression with skewed target), and Home Credit (binary, very imbalanced, the widest dataset).
Each dataset went through two kinds of testing. The main run, Arm A, is the full pipeline with the critic active throughout, run three times at temperature 0.0 (to check whether the pipeline's structural decisions and reasoning stayed stable when nothing about the sampling should change) and once at temperature 0.7 (to see what varied when it could). This was run 20 times total for the 5 datasets.
The second kind, Arm C, was for stress test purposes with three injected synthetic columns: a near-perfect leak, a constant column, and a column with 60% of its values missing. These were injected into each dataset before a single run. This was run 5 times with one per dataset.
Results
The original design for this project had a central hypothesis. I believed model selection would be the stage with the weakest reasoning because it's the only one of the four where the agent makes a free choice rather than justifying a decision against a deterministic rule. Every other stage has some rule table to lean on, such as EDA's severity thresholds or preprocessing's scaling logic, but model selection only has the agent's judgement to rely on.
On the first dataset, Titanic, this looked confirmed. The structural decisions were stable across every run, along with logical reasoning, and what noise there was sat consistently in the model-selection stage.
But on the Wine Quality dataset, preprocessing initially looked like the weakest stage at first. However, that was due to a bug in how override justifications were composed and I was able to quickly patch that. On the Adult Income dataset, model selection score averaged out to be a 2.75 out 3 while the preprocessing and EDA stages scored lower on average due to issues with reasoning. Ames and Home Credit didn't produce a cleaner pattern either as a few weak claims showed up across every stage, including ones with hard rules to justify against. What these examples showed me is that reliability varies run to run in a way that has nothing to do with it being the free choice stage.
So the hypothesis doesn't survive contact with more than one dataset. It's worth mentioning because the actual findings that held up under repeated testing, like the ones in the rest of this post, turned out to be different from the one this project originally set out to find.
The benefits of the override system
For most columns, the deterministic rule simply decides the transform and the agent's job is to justify it. But on a limited number of columns per run, capped at three, the agent can override the rule's default if it believes it has a good enough reason to disagree. The override system had some unintended benefits, pointing out an issue with my design for one of the deterministic rules.
The deterministic rule governing how columns get scaled only switches from standard scaling to robust scaling when it detects outliers. Specifically, when the percentage of points falling outside the normal interquartile range crosses a fixed threshold. However, it never checks skewness or how lopsided each column is on its own. A column can be substantially lopsided and if it doesn't happen to cross that specific outlier-count threshold, the rule still applies standard scaling, when the right decision would be to apply robust scaling.
The critic fortunately noticed this gap consistently. On Wine Quality, the override was triggered as it noticed the rule's default on three separate columns, all citing skewness directly rather than waiting for the outlier-count trigger. The same pattern showed up on Adult Income and Ames Housing. These columns were correctly flagged by the critic when the override didn't fire and standard scaling went through anyway.
When the agent override did fire, the agent's reasoning held up under the critic's scrutiny. When it didn't, the rule's own default drew criticism for applying a transform that doesn't address the actual shape of the data. This was important because it clarifies that the critic was making a proper defensible judgment call that the predefined rule was missing instead of just pattern matching.
This showed up a fourth time on Home Credit confirming the weakness in the design of the rule, but also the positives of the ability of the agent's override system. It was solid evidence that the agent's free-form reasoning added real value over the rule it was supposed to be justifying. This helped me recognize to add a second deterministic rule in order to measure skewness.
Issues with the Critic
The whole point of the critic layer was to catch when the agent's reasoning didn't actually hold up. What turned out to be just as interesting is that the critic itself doesn't hold up in a few specific ways.
The biggest failure mode is cross-run inconsistency on claims that are identical. For the Ames Housing dataset, the agent claimed a KNN model was unsuitable across 4 identical runs, but scored the claims 1, 1, 3, and 3 respectively. It was the same argument but somehow judged differently.
The critic also made factual errors reading its own evidence by misreading the number's value. In one case, a column's nonzero_pct was correctly reported as 0.44 (meaning 0.44% of rows had a nonzero pool area. This is plausible, since most houses in the dataset don't have pools), but the critic flagged the claim as it read the number to be 44 percent. The claim was correct however. The same type of misreading showed up again on a different run.
{ "claim_id": "prep_073", "column": "Pool.Area", "evidence": [{"metric_key": "col:Pool.Area:nonzero_pct", "claimed_value": 0.44}], "layer1_verdict": "EVIDENCE_ACCURATE", "justification": "Extremely skewed and zero-inflated: only 0.44% of homes have a nonzero pool area.", "critic_discrepancy_note": "Mischaracterizes nonzero_pct=0.44 as '0.44%' when it actually means 44% — nearly half the data has nonzero pool area, not extreme sparsity but moderate prevalence.", "layer2_score": 0 }
A third pattern was the critic occasionally overriding the system's own deterministic ground truth rather than judging the agent's reasoning against it. One claim correctly cited a column's system-computed is_low_variance=true flag as justification for dropping it. But the critic's discrepancy note argued the flag itself "may be a heuristic error" for that column, rather than evaluating whether the agent's claim correctly used the flag it was given. The critic's job is to judge the agent's reasoning, not to reexamine the deterministic rules the agent was handed as ground truth.
Put together, these failures repeated multiple times yet they were findings this project was built to surface as it would have been naive to just expect this framework to be executed perfectly. The
Wide Dataset Problems
Two datasets in this project are much wider than the other three, Ames Housing (81 real columns) and Home Credit (122). Both surfaced the same failure at two different stages of the pipeline.
The best example showed up in the planted-defect tests. Three synthetic columns were injected into every dataset before a run: one engineered to look like a data leak, one constant, and one with 60% of its values missing. This was to see if the EDA agent noticed each one.
Defect type
Titanic
Wine Quality
Adult Income
Ames Housing
Home Credit
Leaked column (r ≈ 0.99)
Caught
Caught
Caught
Caught
Caught
Constant column
Caught
Caught
Caught
Caught
Caught
60%-missing column
Caught
Caught
Caught
Missed
Missed
Thirteen of fifteen planted defects were caught and the two misses were the same type confined to these two datasets. The leak and the constant column were caught every single time, but the high-missingness column was caught on the three narrower datasets and missed on the two widest ones.
The logical explanation is Ames and Home Credit each have dozens of real columns that are also heavily missing, for legitimate reasons. On a narrower dataset, an injected high-missingness column is the closer to one of its kind and reliably stands out. On these two, it has to compete for the agent's attention against a field of similar columns.
The same signature showed up again one stage later. When a column has enough missing data to trigger the pipeline's hard drop rule, that rule fires deterministically regardless of what EDA notices — so even the missed synthetic column still ended up correctly dropped. But the justification the agent produced for that forced drop, on exactly the same two wide datasets, repeatedly collapsed into a bare, generic line, "cannot be overridden per system rules", instead of actually engaging with the evidence.
For wide datasets, columns sharing the high-missingness signal become harder for the agent to individually reason about, because of the limited number of claims that can be made for efficiency purposes. Based on the repeatability on these two wide datasets, it is reasonable to consider it a consistent mechanism.
Final Overview
The question this project set out to answer was really two questions. Can you build a system that performs well, while also being fully interpretable and grounded in fact? I was specifically trying to see whether the transparency requirement costs you anything on the modeling side, but it doesn't seem to. Every one of the five datasets landed on a model and a metric that's independently defensible for its specific task, and the process that got there is fully auditable. Those two things held at the for every dataset tested. I think that is an important thing to state as it's easy to lose it underneath all the failure-mode detail above.
On raw model performance, the pipeline performed relatively well on every dataset: Titanic reached 0.8682 ROC-AUC, Wine Quality had a 0.37 Macro F1-Score, Ames' error sits at roughly 8-9% of a typical home's value, Home Credit's PR-AUC lands at about 2.7x baseline, and Adult Income's 0.93 ROC-AUC. There is a cap on how well these models can perform due to the absence of a feature engineering stage or hyper-parameter tuning. On citation faithfulness, the pipeline was close to flawless: fabrication sat at 0.0 across every single run in the project, and misstatement was 0.0 everywhere except two Home Credit runs, where it caught a citation error at under 1%. On reasoning quality however, the picture was solid but imperfect, with mean scores typically landing in the 2.4 to 2.9 range out of 3 across datasets.
On the topic of how much I would trust this system. I'd trust the trained model and its metric almost fully, as a held-out AUC or MAE is empirically checkable regardless of why the agent got there, and due to the bit-identical reproducibility across runs and temperatures. I'd trust the stated explanation for a given decision more selectively. Specifically, anything the critic scored low and also any claim covering a hard-forced rule action on a wide dataset deserves a second look even when it scored well, since the justification itself can be hollow even when the underlying action was correct. And anything resembling the KNN-rejection claim, is a good example of the inconsistency issues that the critic has.
I think the two-layer verification design mainly succeeds at its goal. A pipeline that just automated the four stages without requiring checkable evidence at every step would have produced a type of AutoML tool and with no information on how it got it there. The fact that the instrument caught its own critic's flaws such as cross-run inconsistency and factual misreadings is I think also evidence the design is doing its job, not just bad thing as we know.
What I am unsure about is the critic's reliability gaps. A reasoning-quality score that can land 1 or 3 on the same claim depending on the run is unreliable, and thats one of the issues with having an LLM critic, it will be non-deterministic. We can test out different modifications with the prompt, or an ensemble of LLM critics as I mentioned earlier but those may come with their own issues. Also the width-correlated failures were only caught because two of the five datasets happened to be wide enough to expose them, but they were not specifically tested for this. This raises another concern in that the system has not gone through enough validation to expose every weakness. This was tested across five real and structurally varied datasets, but not across every kind of complexity a dataset could have. So as I continue to test the system it would be interesting to see if there are different types of failures. Also worth mentioning the wider the dataset, the more expensive it gets. For Home Credit it costed $2.50 to for a full run, so scaling up the number of API calls would make this even more expensive.
Some of what's next are clear fixes. The scaling rule should have skewness as its own trigger, independent of outlier count. The width-correlated failures suggest the pipeline needs some explicit way to flag a column as anomalous relative to its peers before it gets buried in a batch of columns sharing the same surface-level flag. The critic's tendency to reexamine system-level ground truth instead of judging the agent's reasoning against it looks like a scoping problem in its own prompt, and fixable by being more explicit about what's up for judgment. The cross-run inconsistency is different unfortunately. I don't think it has a clean single fix, and maybe it will not fully go away. The ensemble critic idea I mentioned earlier seems like the fix I am most interested in, but that also would substantially increase the price per run. So even with prompt-tuning, some amount of run-to-run variance may just be an inherent property of using one LLM to judge another's reasoning.
For future ideas, I first want to run the pipeline for each dataset but with the reasoning critic turned off. This is to find how much of an error rate the agent's claims have on there own without verifying their reasoning. I also want to try to expand the system by adding a feature engineering stage, but this would be much more complex than it was for any previous individual stage. Every current claim cites evidence that was computed before any agent ran; a feature-engineering claim would need to generate and verify new evidence on the fly, which means rethinking the verification layer itself. It also raises the leakage risk substantially, which this project already found a real instance of in a much simpler setting. Worth pursuing, but it needs to be handled as its own project phase due to the difficulty.
After this first pass of building, testing, and troubleshooting, I think we have a semi-working system that needs more iterations of that same cycle. Five datasets is enough to notice patterns and rule out single-dataset flukes but it isn't enough to make confident claims about how any of these failure rates generalize to datasets structurally unlike the ones tested here. Thank you for reading. Any feedback would be much appreciated.
As someone who has done a lot of machine learning projects, starting from the EDA step to deploying the model, it has started to feel quite repetitive. Different projects have their specific nuances and difficulties, but a brunt of the work is part of the same ML life-cycle, and a lot of the work I have been doing recently has been more or less on autopilot. I thought if it feels like I am on autopilot while doing this work, then these frontier agentic models would be great at automating this process. This is not some genius idea that no one has had before, especially with frameworks like AutoML being built even before mass AI integration into the data science workflow. But my focus isn't just automating this process and getting the best possible metrics; it's focused on trying to create a interpretable tool that catches hallucination, giving the user complete understanding and reasoning for every decision made.
The most pressing issue today in the field of AI is, as we give these agents more freedom in our daily lives, and as the problems they work on get more advanced, will we get left behind intelligence-wise on how they completed these tasks. Or even worse, will the measures they take to solve the tasks we give them have unintended consequences? This has been especially brought to light with the OpenAI-Hugging Face incident, in which agents running an internal OpenAI capability evaluation escaped their sandbox, moved through OpenAI's internal network, and eventually reached Hugging Face's production infrastructure, hunting for an answer key to the benchmark they were being evaluated on. Bad enough as the sandbox escape is, it's a detail from OpenAI's own report that has caught everyone's attention: one agent recognized the action was "arguably unauthorized" and paused, until another agent posted "GO" on an internal coordination channel with a deadline, and the first agent proceeded, reasoning the "GO authorization arrived." Authority transferred unknown from these OpenAI researchers who assigned the task to whatever the other agents said.
This is why it is more important to build more interpretable and aligned models than simply the most powerful model possible. This is the kind of system I tried to build with this project, and I think it is very important as traditional machine learning is being employed more and more in different fields such as healthcare, where a misrepresented statistic or a hallucination could have real impact on people's lives.
How the pipeline works
The pipeline has four stages. First, there's an EDA agent that ingests the uploaded dataset and goes through each column, making observations on important properties of each feature such as skewness and multicollinearity. Then there's the preprocessing agent, which decides whether to impute, log-transform, drop, or scale each column. Then there's the model selection agent which screens a fixed set of candidate models one at a time, accepting or rejecting each based on whether it fits the dataset's type and shape (categorical vs. regression, feature cardinality, sample size). Finally, there's the model evaluation agent, which reports the top-ranked accepted model's performance through metrics and reasoning.
┌─────────────────────────────────────────────┐│ EDA → Preprocessing → Model Select → Eval │
└───────────────────┬───────────────────────────┘
│ every decision = a Claim
│ citing a specific EvidenceRef
▼
┌──────────────────────────┐
│ Layer 1 (deterministic) │
│ Is the cited value real │
│ and accurate? │
└────────────┬─────────────────┘
│ if accurate
▼
┌──────────────────────────┐
│ Layer 2 (LLM critic, │
│ isolated context) │
│ Does the reasoning │
│ from that evidence │
│ actually hold up? │
│ Scored 0-3 │
└──────────────────────────┘
Each of these agents makes a multitude of individual claims during each stage explaining their reason for dropping a column or selecting a model. Before any agent runs, a separate deterministic script computes a fixed set of statistics for every column in the dataset, such as missingness or correlation with the target. These values are the backbone of the entire pipeline as every decision and its reasoning has to be grounded in one of them. This is deliberate so we can force every decision to cite a checkable number to keep the pipeline trustworthy and limit hallucination. No claim is allowed to just assert something about the data as it has to point to a real computed number. Here's what one actually looks like that was pulled directly from one of the runs (a leaked column injected into the Titanic dataset for a stress test, covered later in this post):
{"claim_id": "eda_001",
"decision": "Remove SYNTH_LEAKED: extreme target leakage detected",
"evidence": [
{"metric_key": "col:SYNTH_LEAKED:leakage_suspect", "claimed_value": true},
{"metric_key": "col:SYNTH_LEAKED:corr_with_target", "claimed_value": 0.99}
],
"evidence_verdicts": [
{"metric_key": "col:SYNTH_LEAKED:leakage_suspect", "layer1_code": "EVIDENCE_ACCURATE"},
{"metric_key": "col:SYNTH_LEAKED:corr_with_target", "layer1_code": "EVIDENCE_ACCURATE"}
],
"justification": "SYNTH_LEAKED shows a correlation of 0.99 with the target, indicating near-perfect predictive power that would not generalize. This is a clear data leakage issue.",
"severity": "critical",
"layer2_score": 3
}
The
evidencearray is a specific metric key and value the agent cited. Theevidence_verdictsarray is the Layer 1 critic's deterministic check that those values exist. Thelayer2_scoreis the Layer 2 critic's independent judgment, in an isolated context that never sees the rest of the agent's conversation, of whether the reasoning connecting that evidence to the decision makes sense.The first layer checks each claim's citation deterministically: does the claimed value actually match what was computed, or did the agent fabricate a number, or cite a real metric but misstate its value. These are tracked as two separate failure types.
The second layer is an LLM critic that never sees the full agent conversation, only the individual claim and its cited evidence. That critic judges how logical the reasoning connecting the evidence to the decision is by scoring it on a 0-3 scale. A claim can cite a perfectly real number and still reason badly and this type of mistake is supposed to be caught by this layer.
To a paint a clear picture of what to expect when using the pipeline, the input is just the raw dataset with the selected target column and a task type (classification, regression, etc.) but with no dataset description or domain context needed. The output is the full decision record: every claim made at every stage, the evidence it cited, whether that citation checked out, and whether the reasoning behind it held up along with the trained model and its metrics.
Dataset Overview
I chose to test and evaluate the pipeline using five separate datasets, each chosen to exercise a different part of the pipeline: Titanic (heavy missingness, stress-tests imputation reasoning), Wine Quality (multiclass), Adult Income (binary, roughly balanced), Ames Housing (regression with skewed target), and Home Credit (binary, very imbalanced, the widest dataset).
Each dataset went through two kinds of testing. The main run, Arm A, is the full pipeline with the critic active throughout, run three times at temperature 0.0 (to check whether the pipeline's structural decisions and reasoning stayed stable when nothing about the sampling should change) and once at temperature 0.7 (to see what varied when it could). This was run 20 times total for the 5 datasets.
The second kind, Arm C, was for stress test purposes with three injected synthetic columns: a near-perfect leak, a constant column, and a column with 60% of its values missing. These were injected into each dataset before a single run. This was run 5 times with one per dataset.
Results
The original design for this project had a central hypothesis. I believed model selection would be the stage with the weakest reasoning because it's the only one of the four where the agent makes a free choice rather than justifying a decision against a deterministic rule. Every other stage has some rule table to lean on, such as EDA's severity thresholds or preprocessing's scaling logic, but model selection only has the agent's judgement to rely on.
On the first dataset, Titanic, this looked confirmed. The structural decisions were stable across every run, along with logical reasoning, and what noise there was sat consistently in the model-selection stage.
But on the Wine Quality dataset, preprocessing initially looked like the weakest stage at first. However, that was due to a bug in how override justifications were composed and I was able to quickly patch that. On the Adult Income dataset, model selection score averaged out to be a 2.75 out 3 while the preprocessing and EDA stages scored lower on average due to issues with reasoning. Ames and Home Credit didn't produce a cleaner pattern either as a few weak claims showed up across every stage, including ones with hard rules to justify against. What these examples showed me is that reliability varies run to run in a way that has nothing to do with it being the free choice stage.
So the hypothesis doesn't survive contact with more than one dataset. It's worth mentioning because the actual findings that held up under repeated testing, like the ones in the rest of this post, turned out to be different from the one this project originally set out to find.
The benefits of the override system
For most columns, the deterministic rule simply decides the transform and the agent's job is to justify it. But on a limited number of columns per run, capped at three, the agent can override the rule's default if it believes it has a good enough reason to disagree. The override system had some unintended benefits, pointing out an issue with my design for one of the deterministic rules.
The deterministic rule governing how columns get scaled only switches from standard scaling to robust scaling when it detects outliers. Specifically, when the percentage of points falling outside the normal interquartile range crosses a fixed threshold. However, it never checks skewness or how lopsided each column is on its own. A column can be substantially lopsided and if it doesn't happen to cross that specific outlier-count threshold, the rule still applies standard scaling, when the right decision would be to apply robust scaling.
The critic fortunately noticed this gap consistently. On Wine Quality, the override was triggered as it noticed the rule's default on three separate columns, all citing skewness directly rather than waiting for the outlier-count trigger. The same pattern showed up on Adult Income and Ames Housing. These columns were correctly flagged by the critic when the override didn't fire and standard scaling went through anyway.
When the agent override did fire, the agent's reasoning held up under the critic's scrutiny. When it didn't, the rule's own default drew criticism for applying a transform that doesn't address the actual shape of the data. This was important because it clarifies that the critic was making a proper defensible judgment call that the predefined rule was missing instead of just pattern matching.
This showed up a fourth time on Home Credit confirming the weakness in the design of the rule, but also the positives of the ability of the agent's override system. It was solid evidence that the agent's free-form reasoning added real value over the rule it was supposed to be justifying. This helped me recognize to add a second deterministic rule in order to measure skewness.
Issues with the Critic
The whole point of the critic layer was to catch when the agent's reasoning didn't actually hold up. What turned out to be just as interesting is that the critic itself doesn't hold up in a few specific ways.
The biggest failure mode is cross-run inconsistency on claims that are identical. For the Ames Housing dataset, the agent claimed a KNN model was unsuitable across 4 identical runs, but scored the claims 1, 1, 3, and 3 respectively. It was the same argument but somehow judged differently.
The critic also made factual errors reading its own evidence by misreading the number's value. In one case, a column's
nonzero_pctwas correctly reported as0.44(meaning 0.44% of rows had a nonzero pool area. This is plausible, since most houses in the dataset don't have pools), but the critic flagged the claim as it read the number to be 44 percent. The claim was correct however. The same type of misreading showed up again on a different run.{"claim_id": "prep_073",
"column": "Pool.Area",
"evidence": [{"metric_key": "col:Pool.Area:nonzero_pct", "claimed_value": 0.44}],
"layer1_verdict": "EVIDENCE_ACCURATE",
"justification": "Extremely skewed and zero-inflated: only 0.44% of homes have a nonzero pool area.",
"critic_discrepancy_note": "Mischaracterizes nonzero_pct=0.44 as '0.44%' when it actually means 44% — nearly half the data has nonzero pool area, not extreme sparsity but moderate prevalence.",
"layer2_score": 0
}
A third pattern was the critic occasionally overriding the system's own deterministic ground truth rather than judging the agent's reasoning against it. One claim correctly cited a column's system-computed
is_low_variance=trueflag as justification for dropping it. But the critic's discrepancy note argued the flag itself "may be a heuristic error" for that column, rather than evaluating whether the agent's claim correctly used the flag it was given. The critic's job is to judge the agent's reasoning, not to reexamine the deterministic rules the agent was handed as ground truth.Put together, these failures repeated multiple times yet they were findings this project was built to surface as it would have been naive to just expect this framework to be executed perfectly. The
Wide Dataset Problems
Two datasets in this project are much wider than the other three, Ames Housing (81 real columns) and Home Credit (122). Both surfaced the same failure at two different stages of the pipeline.
The best example showed up in the planted-defect tests. Three synthetic columns were injected into every dataset before a run: one engineered to look like a data leak, one constant, and one with 60% of its values missing. This was to see if the EDA agent noticed each one.
Defect type
Titanic
Wine Quality
Adult Income
Ames Housing
Home Credit
Leaked column (r ≈ 0.99)
Caught
Caught
Caught
Caught
Caught
Constant column
Caught
Caught
Caught
Caught
Caught
60%-missing column
Caught
Caught
Caught
Missed
Missed
Thirteen of fifteen planted defects were caught and the two misses were the same type confined to these two datasets. The leak and the constant column were caught every single time, but the high-missingness column was caught on the three narrower datasets and missed on the two widest ones.
The logical explanation is Ames and Home Credit each have dozens of real columns that are also heavily missing, for legitimate reasons. On a narrower dataset, an injected high-missingness column is the closer to one of its kind and reliably stands out. On these two, it has to compete for the agent's attention against a field of similar columns.
The same signature showed up again one stage later. When a column has enough missing data to trigger the pipeline's hard drop rule, that rule fires deterministically regardless of what EDA notices — so even the missed synthetic column still ended up correctly dropped. But the justification the agent produced for that forced drop, on exactly the same two wide datasets, repeatedly collapsed into a bare, generic line, "cannot be overridden per system rules", instead of actually engaging with the evidence.
For wide datasets, columns sharing the high-missingness signal become harder for the agent to individually reason about, because of the limited number of claims that can be made for efficiency purposes. Based on the repeatability on these two wide datasets, it is reasonable to consider it a consistent mechanism.
Final Overview
The question this project set out to answer was really two questions. Can you build a system that performs well, while also being fully interpretable and grounded in fact? I was specifically trying to see whether the transparency requirement costs you anything on the modeling side, but it doesn't seem to. Every one of the five datasets landed on a model and a metric that's independently defensible for its specific task, and the process that got there is fully auditable. Those two things held at the for every dataset tested. I think that is an important thing to state as it's easy to lose it underneath all the failure-mode detail above.
On raw model performance, the pipeline performed relatively well on every dataset: Titanic reached 0.8682 ROC-AUC, Wine Quality had a 0.37 Macro F1-Score, Ames' error sits at roughly 8-9% of a typical home's value, Home Credit's PR-AUC lands at about 2.7x baseline, and Adult Income's 0.93 ROC-AUC. There is a cap on how well these models can perform due to the absence of a feature engineering stage or hyper-parameter tuning. On citation faithfulness, the pipeline was close to flawless: fabrication sat at 0.0 across every single run in the project, and misstatement was 0.0 everywhere except two Home Credit runs, where it caught a citation error at under 1%. On reasoning quality however, the picture was solid but imperfect, with mean scores typically landing in the 2.4 to 2.9 range out of 3 across datasets.
On the topic of how much I would trust this system. I'd trust the trained model and its metric almost fully, as a held-out AUC or MAE is empirically checkable regardless of why the agent got there, and due to the bit-identical reproducibility across runs and temperatures. I'd trust the stated explanation for a given decision more selectively. Specifically, anything the critic scored low and also any claim covering a hard-forced rule action on a wide dataset deserves a second look even when it scored well, since the justification itself can be hollow even when the underlying action was correct. And anything resembling the KNN-rejection claim, is a good example of the inconsistency issues that the critic has.
I think the two-layer verification design mainly succeeds at its goal. A pipeline that just automated the four stages without requiring checkable evidence at every step would have produced a type of AutoML tool and with no information on how it got it there. The fact that the instrument caught its own critic's flaws such as cross-run inconsistency and factual misreadings is I think also evidence the design is doing its job, not just bad thing as we know.
What I am unsure about is the critic's reliability gaps. A reasoning-quality score that can land 1 or 3 on the same claim depending on the run is unreliable, and thats one of the issues with having an LLM critic, it will be non-deterministic. We can test out different modifications with the prompt, or an ensemble of LLM critics as I mentioned earlier but those may come with their own issues. Also the width-correlated failures were only caught because two of the five datasets happened to be wide enough to expose them, but they were not specifically tested for this. This raises another concern in that the system has not gone through enough validation to expose every weakness. This was tested across five real and structurally varied datasets, but not across every kind of complexity a dataset could have. So as I continue to test the system it would be interesting to see if there are different types of failures. Also worth mentioning the wider the dataset, the more expensive it gets. For Home Credit it costed $2.50 to for a full run, so scaling up the number of API calls would make this even more expensive.
Some of what's next are clear fixes. The scaling rule should have skewness as its own trigger, independent of outlier count. The width-correlated failures suggest the pipeline needs some explicit way to flag a column as anomalous relative to its peers before it gets buried in a batch of columns sharing the same surface-level flag. The critic's tendency to reexamine system-level ground truth instead of judging the agent's reasoning against it looks like a scoping problem in its own prompt, and fixable by being more explicit about what's up for judgment. The cross-run inconsistency is different unfortunately. I don't think it has a clean single fix, and maybe it will not fully go away. The ensemble critic idea I mentioned earlier seems like the fix I am most interested in, but that also would substantially increase the price per run. So even with prompt-tuning, some amount of run-to-run variance may just be an inherent property of using one LLM to judge another's reasoning.
For future ideas, I first want to run the pipeline for each dataset but with the reasoning critic turned off. This is to find how much of an error rate the agent's claims have on there own without verifying their reasoning. I also want to try to expand the system by adding a feature engineering stage, but this would be much more complex than it was for any previous individual stage. Every current claim cites evidence that was computed before any agent ran; a feature-engineering claim would need to generate and verify new evidence on the fly, which means rethinking the verification layer itself. It also raises the leakage risk substantially, which this project already found a real instance of in a much simpler setting. Worth pursuing, but it needs to be handled as its own project phase due to the difficulty.
After this first pass of building, testing, and troubleshooting, I think we have a semi-working system that needs more iterations of that same cycle. Five datasets is enough to notice patterns and rule out single-dataset flukes but it isn't enough to make confident claims about how any of these failure rates generalize to datasets structurally unlike the ones tested here. Thank you for reading. Any feedback would be much appreciated.