TL;DR a few model providers are learning from private conversations across many organisations, giving their models better representations of how those organisations secure their systems and plausibly making semi-autonomous attacks easier.
Private credentials are given directly to LLMs at scale. Beyond the same WiFi password being reused across office-wifis or stored in a Google doc, people post API keys straight into prompts, hardcode them in files, copy them into chats on personal accounts, etc.
I’m not sure where chat data enters model development and so what mitigations exist for the inclusion of credentials. I imagine chat data could used as part of continual pre-training, mid-training, or via selected examples used in post-training for supervised fine-tuning or RLHF - several points of intervention where credentials could be removed. In any case, development delays make the risk caused by recalling live credentials more uncertain. Public data already leaks credentials at scale, while GPT-4 and GPT-4o generate guesses that look plausible but don’t closely match how real users build passwords.
Beyond leveraging previously leaked credentials, or that provided by continual training on chat data, the concern of this post is that even with credentials stripped, models are building better representations of how real organisations work internally, and how people do their jobs in them. Models seem pretty good at inferring things about people, which isn’t prevented by anonymisation, or in this case scrubbing credentials, so it is plausible that representations of people-in-jobs would be learn. Aggregated across companies, ordinary details about people's practices within their roles, and those of the organization, might become useful priors about where similar organisations are weak.
I’d guess that some weak practices are already well understood by models (like unsecured fish tank thermometers allowing casino database access). The additional danger is about predicting and situating these weaknesses in organisational context, allowing more consistent, directed guesses - and plausibly the kind of semi-autonomous attacks we’re seeing.
Is fishtank thermometer security knowledge sufficient to execute an exploit? What about to think about and plan one?
Hence: are models trained on cross-company chat data becoming better at inferring organisational vulnerabilities?
Open-source models trained on public data lack this particular advantage, unless a large host collects and trains on their users’ chats - so plausibly this mainly happens in a world where there are few frontier labs, and they aggregate across verticals. There might be a slow takeoff dynamic across model generations. Wider adoption produces a broader cross-company corpus, which produces more useful models and further adoption. Opsec failures become increasingly predictable as model capabilities and user chat data integration scale, allowing agents to generate reasonable next-attack options against unfamiliar targets. I could see a catastrophic scenario here where agents can gain access to critical infrastructure relatively easily.
A better understanding of exactly how next targets/strategy is performed by the hacker might change my mind, for example if this shows that choosing the next attack isn't so much of a bottleneck, or information comes from multiple sources (not just the fishtank).
Without further analysis of the actual representations conferred by chat data, evaluating if chat data is responsible for better "hacking" abilities seems a hard question to test (though worthwhile). It requires differentiating general capability scaling with capability scaling along specific axes (a problem partially solved by evals), and linking that to specific training data (a problem not yet solved by mechanistic interpretability, as far as I know). It's not even clear that we can link specific representations (as they evolve) to specific outputs. But anyway, that's the next casino's problem.
"Thomas Schafer aquarelle" + "Evolution of modularity"
Private credentials are given directly to LLMs at scale. Beyond the same WiFi password being reused across office-wifis or stored in a Google doc, people post API keys straight into prompts, hardcode them in files, copy them into chats on personal accounts, etc.
We know that since mid 2026 Anthropic may use Free, Pro and Max chats to train future models unless users opt out. OpenAI has a similar policy for consumer ChatGPT. Both exclude business and API products by default, but employees almost certainly also use personal accounts for work.
I’m not sure where chat data enters model development and so what mitigations exist for the inclusion of credentials. I imagine chat data could used as part of continual pre-training, mid-training, or via selected examples used in post-training for supervised fine-tuning or RLHF - several points of intervention where credentials could be removed. In any case, development delays make the risk caused by recalling live credentials more uncertain. Public data already leaks credentials at scale, while GPT-4 and GPT-4o generate guesses that look plausible but don’t closely match how real users build passwords.
Beyond leveraging previously leaked credentials, or that provided by continual training on chat data, the concern of this post is that even with credentials stripped, models are building better representations of how real organisations work internally, and how people do their jobs in them. Models seem pretty good at inferring things about people, which isn’t prevented by anonymisation, or in this case scrubbing credentials, so it is plausible that representations of people-in-jobs would be learn. Aggregated across companies, ordinary details about people's practices within their roles, and those of the organization, might become useful priors about where similar organisations are weak.
I’d guess that some weak practices are already well understood by models (like unsecured fish tank thermometers allowing casino database access). The additional danger is about predicting and situating these weaknesses in organisational context, allowing more consistent, directed guesses - and plausibly the kind of semi-autonomous attacks we’re seeing.
Is fishtank thermometer security knowledge sufficient to execute an exploit? What about to think about and plan one?
A recent attack seems to have been largely model-executed across around thirty targets. This may be skilled attackers using the model’s general software and networking capabilities. Better organisational representations could also help the model choose reasonable next actions.
Hence: are models trained on cross-company chat data becoming better at inferring organisational vulnerabilities?
Open-source models trained on public data lack this particular advantage, unless a large host collects and trains on their users’ chats - so plausibly this mainly happens in a world where there are few frontier labs, and they aggregate across verticals. There might be a slow takeoff dynamic across model generations. Wider adoption produces a broader cross-company corpus, which produces more useful models and further adoption. Opsec failures become increasingly predictable as model capabilities and user chat data integration scale, allowing agents to generate reasonable next-attack options against unfamiliar targets. I could see a catastrophic scenario here where agents can gain access to critical infrastructure relatively easily.
A better understanding of exactly how next targets/strategy is performed by the hacker might change my mind, for example if this shows that choosing the next attack isn't so much of a bottleneck, or information comes from multiple sources (not just the fishtank).
Without further analysis of the actual representations conferred by chat data, evaluating if chat data is responsible for better "hacking" abilities seems a hard question to test (though worthwhile). It requires differentiating general capability scaling with capability scaling along specific axes (a problem partially solved by evals), and linking that to specific training data (a problem not yet solved by mechanistic interpretability, as far as I know). It's not even clear that we can link specific representations (as they evolve) to specific outputs. But anyway, that's the next casino's problem.
"Thomas Schafer aquarelle" + "Evolution of modularity"