TL;DR We think that collecting data generated as part of the AI safety research process can be used to train models to be better safety researchers. In our SPAR Fall 2026 project, mentored by Alec Harris, we’re piloting work to collect this data from organizations and independent researchers and store it in a centralized repository.
We’re sharing this post for awareness and to gather feedback. We are also looking to find researchers or organizations to gather this data from.
Types of Data
A bottleneck to safety research in AI is strong in-domain conceptual reasoning that isn’t represented well in the general corpus of reasoning data. We think that artifacts in the safety domain that demonstrate the deliberation steps before an output, where ideas are modified and directions are chosen, can make AI differentially better at safety research, i.e. improving safety research ability without just improving general capabilities. This can be used to teach taste and the process of research design, not just what a good output looks like. We’re first prioritizing these types of data that might demonstrate this for research projects:
Documents (e.g. google docs or latex docs) with comments and editing revision history
Meeting transcripts
Emails and messages
Screen recordings of research plans being reviewed
Currently, our philosophy for this project is to optimize for volume and coverage of collected data over focusing on collecting data for a particular training pipeline, as we expect infrastructure for aggregating data to be the bottleneck.
We are building requirements around consenting and opt-ins and are studying concerns around privacy, security and other sensitive information.
Providers of Data
We’re envisioning 2 primary types:
Individual researchers - Individual researchers might be able to provide data more freely, but the quality of the data might vary more.
Organizations - Organizations might already have data collection set up internally. There is scope for us to establish a working relationship, for continued data collection at scale rather than reaching out to a lot of individual researchers.
Consumers of Data
By data consumers, we mean parties with access to the data repository that are interested in using the data to build something aimed at improving AI safety research. For example, a safety organization might filter collected document revision histories and meeting transcripts where ideas were debated, and use this data to fine-tune a model to give feedback on safety research ideas. We intend to initially store the raw data collected, with minimal labelling, and only open it up to a few parties trusted by the providers, who can filter the data as they see fit.
We have open questions about who should have access to the data and what controls might need to exist, but given our current capacity, our plan for this project is to focus on collecting and storing the data.
Data Control
We are currently angling to let data providers choose what is collected, and store this in a centralized repository, where our team and the consumers have access.
We would do basic tagging on data, such as labelling the associated data provider, date, data type, and the source of the data.
We are still considering the requirements around consent and opt-ins, but by default:
Providers can request to revoke their data
We would handle storing the data in the centralized repository
Only consumers approved by the providers could access the data
Our idea here is that this would make providers be more comfortable providing data. Our current plan is that all consumers approved by a provider would have access to the same centralized data repository.
We would require consent from any third parties involved in relevant transcripts
In future work, we envision a general centralized database for more wide-spread use of data, as well as private repositories for each provider where they can limit access to sensitive data to particular parties.
Call to Action
How you can help:
If you’re actively working on research projects: We need lead users or organizations for our pilot to test collecting and storing data. Please give us feedback or let us know if you are interested here - the form takes < 5 minutes to complete.
Anyone: Comment here and we can discuss any ideas, questions or concerns.
TL;DR We think that collecting data generated as part of the AI safety research process can be used to train models to be better safety researchers. In our SPAR Fall 2026 project, mentored by Alec Harris, we’re piloting work to collect this data from organizations and independent researchers and store it in a centralized repository.
We’re sharing this post for awareness and to gather feedback. We are also looking to find researchers or organizations to gather this data from.
Types of Data
A bottleneck to safety research in AI is strong in-domain conceptual reasoning that isn’t represented well in the general corpus of reasoning data. We think that artifacts in the safety domain that demonstrate the deliberation steps before an output, where ideas are modified and directions are chosen, can make AI differentially better at safety research, i.e. improving safety research ability without just improving general capabilities. This can be used to teach taste and the process of research design, not just what a good output looks like. We’re first prioritizing these types of data that might demonstrate this for research projects:
Currently, our philosophy for this project is to optimize for volume and coverage of collected data over focusing on collecting data for a particular training pipeline, as we expect infrastructure for aggregating data to be the bottleneck.
We are building requirements around consenting and opt-ins and are studying concerns around privacy, security and other sensitive information.
Providers of Data
We’re envisioning 2 primary types:
Consumers of Data
By data consumers, we mean parties with access to the data repository that are interested in using the data to build something aimed at improving AI safety research. For example, a safety organization might filter collected document revision histories and meeting transcripts where ideas were debated, and use this data to fine-tune a model to give feedback on safety research ideas. We intend to initially store the raw data collected, with minimal labelling, and only open it up to a few parties trusted by the providers, who can filter the data as they see fit.
We have open questions about who should have access to the data and what controls might need to exist, but given our current capacity, our plan for this project is to focus on collecting and storing the data.
Data Control
We are currently angling to let data providers choose what is collected, and store this in a centralized repository, where our team and the consumers have access.
We would do basic tagging on data, such as labelling the associated data provider, date, data type, and the source of the data.
We are still considering the requirements around consent and opt-ins, but by default:
In future work, we envision a general centralized database for more wide-spread use of data, as well as private repositories for each provider where they can limit access to sensitive data to particular parties.
Call to Action
How you can help:
If you’re actively working on research projects: We need lead users or organizations for our pilot to test collecting and storing data. Please give us feedback or let us know if you are interested here - the form takes < 5 minutes to complete.
Anyone: Comment here and we can discuss any ideas, questions or concerns.