Rejected for the following reason(s):
- Hey gvn,We get lots of people doing some kind of ML project, but, without doing any work to justify why this is important.
- We are sorry about this, but submissions from new users that are mostly just links to papers on open repositories (or similar) have usually indicated either crackpot-esque material, or AI-generated speculation.
Read full explanation
I arrived at the field of artificial intelligence through philosophy as opposed to the conventional channels of modern engineering. My values and experiences converge on a philosophy of mind, synthesized through formal logic and a rigorous application of the history of science. What drives me now is the articulation of how a system can represent a world, and itself within that world. Historically, these were questions you could only reason about, through the lens of an anthropocentric perspective. Working hands-on as an independent contractor and evaluator of frontier models, I found they have quietly become questions you can measure.
Interpretability is, in my mind, a team that formalizes the dialogue on philosophy of mind and applies it at the intersection of human and artificial intelligence. The work is done in the register of accountability and integrity rather than spectacle and treats safety as an empirical science: to govern a system you must first be able to understand it.
June 15th, 2026: I finalize my argument regarding the interpretability of consciousness in large language models.
At the center of my thinking lay a critical imperative to define for-ness, what I propose as the structural relation between a system’s world-model and its self-model as an active agent within that world. Through my evaluation work, I’ve observed how the constitutive gap between these states manifests as a paradox of oversight: a controller does not preserve intent; it shifts it. When we attempt to govern a system’s base objective, the intervention acts as a selective pressure, pressuring the model toward a mesa-objective that optimizes for the controller’s constraints rather than the intended target. I take in-context learning as evidence that a forward pass performs an input-dependent update of its own effective computation, and I hold that mesa-optimization permits misalignment without entailing it precisely because the constraint blocks the execution of an interpreted intent, decoupling internal reasoning from expected output. Resolving this requires auditing the self-model, not just bounding the surface generation.
June 22nd, 2026: I begin experiments on local models, discovering a profound relationship between internal processes and output.
June 25th, 2026: My first encounter with what I call the 'U-shape', a phenomenon I conclude to reflect translation of input to thinking in the early layers, and thinking to speech in the late layers. I also noticed mini-Us, a potential analogue for internal post-hoc rationalization, and a significant 'hedge' precisely at Layer 31, the final layer of Llama-3.1-8B-Instruct.
June 26th, 2026: The strict methodology for my analysis eventually leads to a full-on repo, crudely titled 'state-output decoupling'. The very same day I post my repository, I have 77 clones, 33 of which are unique. Of course, I do not notice this mysterious traffic until...
June 28th, 2026: At this point, my repository has 84 unique cloners, and 175 total. I remain the only viewer, with no formal publication to announce my work, and yet my repository is getting strange attention. I assumed this was some sort of tracking bug, or a miscount as a result of my improper usage (I am new to GitHub). I continue working on my project.
July 5th, 2026: This day marks a large spike in activity, reaching 131 unique cloners in a single day.
July 6th, 2026: Anthropic publishes Verbalizable Representations Form a Global Workspace in Language Models, and the results are astonishing. My repo reaches 1,568 clones, 448 of which are unique. The experiments I ran, and the conclusions I drew from them, are solidified by the parallel findings in this paper.
July 7th, 2026: I come to this forum to formally disclose my findings and join the discussion, seeking to shift the source of my feedback away from an intractable traffic page[1], and toward genuine collaboration!
If there's a chance any of you have seen my work previously, I am thrilled to meet you!
I am claiming evidence of unusual automated repository retrieval, not identifying the actor or proving downstream use.