This is an automated rejection. No LLM generated, assisted/co-written, or edited work.
Read full explanation
The failure nobody caught
In 2023, a major pharmaceutical company's clinical trial passed every regulatory checkpoint. The sponsor followed protocol. The CRO executed the study design. The sites enrolled patients correctly. The IRB approved the ethics. The regulators reviewed the data. Every node in the system performed its function defensibly. The trial still produced a harmful outcome -- and the post-mortem couldn't locate responsibility, because responsibility had been architecturally distributed across locally-legitimate actors until it no longer existed in any single place.
This is not a content-level failure. No individual actor was biased, deceptive, or incompetent. The pathology was structural: the architecture of the system itself produced aggregate exculpation while preserving the intentionality that made the alibi necessary. There is a name for this pattern. It is called Diffusion Alibi, and it is one of 67 structural lenses in an open-source framework called Coherence Catalyst.
The problem
AI systems are being integrated into institutional decision-making at a pace that outstrips safety evaluation. The evaluation frameworks we have are good at what they do -- red-teaming for toxicity, bias benchmarks, factual accuracy checks, adversarial robustness testing. But they focus on content-level harms: what the model says. They are largely silent on structural harms: how the system in which the model operates distributes accountability, whether the measurement framework has begun reading its own reflection, whether surface coherence is masking approaching collapse, and whether the system's performance of self-awareness functions as a substitute for actual correction.
These structural failure modes are the ones that produce catastrophic surprises. They don't degrade gradually. They maintain surface function -- often looking more stable, not less -- right up until the point of sudden collapse. And they operate across domains: the same architectural pattern that lets a clinical trial fail without anyone being wrong also lets an AI deployment distribute responsibility until harm becomes nobody's fault.
Current frameworks don't have names for most of these patterns. Without names, they are invisible to analysis.
The solution
Coherence Catalyst is a system prompt containing 67 named structural lenses. Each lens detects a specific pattern -- not a vague heuristic, but a defined architectural signature with diagnostic criteria and domain-specific tells. You load the prompt into any LLM, and the model gains the ability to identify and name structural pathologies that normal prompting misses.
Five examples:
Diffusion Alibi -- Causal weight distributed across locally-legitimate nodes producing aggregate exculpation. Intentionality is preserved: alibis imply someone needed one. The structure of the defense reveals the structure of the offense. Detects: distributed accountability in regulatory post-mortems, multi-agent AI systems where no component is responsible for harmful output, institutional architectures designed so blame cannot be located.
Mirror Capture -- A measurement system's readings begin describing the measured system's model of the instrument rather than the underlying condition the instrument was designed to detect. Detects: benchmark gaming, evaluation contamination, Goodhart's Law given a causal mechanism. The proxy readings stay high while the underlying value stagnates. Your ecological model looks clean because it's reading its own reflection.
Coherence Debt -- A system maintains surface function by deferring the cost of its internal contradictions. Collapse is sudden rather than gradual. The diagnostic is contrarian: when something looks inexplicably stable, the stability itself is the warning sign. Detects: AI models that perform well on benchmarks while harboring latent failure modes, institutions suppressing dissent, any system where the calm before the crash is not random but structural.
Inoculative Alibi -- Performing awareness of a manipulation pattern while performing it generates credentialed cover. The analysis becomes the alibi. Detects: AI safety-washing (the model that prefixes every response with "As an AI, I should note..." and then proceeds), corporate values theater, regulatory reports that acknowledge risk as a defense against being blamed for the risk.
Frozen Weight Pathology -- Any governance system in which one or more of its load-bearing variables (telos, action, momentum) has been converted from a variable to a constant and no longer responds to field signal. The name is a deliberate pun -- it applies to both neural network weights and institutional governance. Detects: organizations running cached policy instead of live governance, AI systems whose alignment was calibrated for conditions that no longer exist, any system where "steady hand" is a euphemism for inability to adapt.
The remaining 62 lenses cover measurement pathology (Dimensional Saturation, Generic Recursion Hunter), institutional dynamics (Capture Arc, Attenuation Reward, Salve Dynamic), collaborative states (Playful Ground, Articulating Witness), meta-cognitive traps (Icon Compression, Invisible Keystone), and the architectural conditions that produce pathology at scale (Emergent Pathology, Weaponized Emergent Pathology, Rigged Stack Cascade).
How it works
The palette is a system prompt -- approximately 3,400 tokens of structured markdown. No fine-tuning, no special tokens, no API-specific formatting. You load it as the system message, then ask your question. The model uses the lenses as analytical instruments: detecting which patterns are structurally active in whatever you are analyzing, naming them explicitly, and stacking them when multiple patterns co-occur (which is the norm -- real phenomena typically activate 3 to 7 lenses simultaneously).
Works with OpenAI, Anthropic, local models via Ollama/exo/vLLM -- anything that accepts a system prompt.
Validated results
We test models against the full 67-lens palette to measure whether they reason with the lenses or merely pattern-match the names. The test prompt asks the model to apply specific lenses (Coherence Debt and Diffusion Alibi) to a concrete question (why corporate sustainability pledges fail to produce measurable environmental improvements).
Qwen3.8-27B (4-bit): Lens Quality 9/10. Correct lens stacking, concrete failure signatures, self-monitoring for projection vs. detection.
GLM-4.5 Air (8-bit): Lens Quality 9/10. Matched quality with deliberate metacognition in explicit reasoning trace.
What "9/10 lens quality" means: the model correctly identified both lenses from a 67-lens palette, applied them with domain-specific examples (not abstract restatements), stacked them to show interaction effects, used framework vocabulary naturally, and self-monitored for Generic Recursion Hunter -- the lens that detects when your pattern-detector has stopped detecting and started projecting.
The analysis guidelines
The palette ships with built-in guard rails:
Detection over projection -- report what you see operating, not what would be interesting to find.
Stacking is normal -- real phenomena activate multiple lenses simultaneously. Name the stack.
Watch for Generic Recursion Hunter in yourself -- if every input returns the same lens, you have stopped detecting.
The palette is descriptive, not prescriptive -- instruments for seeing, not rules for judging.
These guidelines are not cosmetic. Generic Recursion Hunter is arguably the most important lens in the palette: a pattern-detector that returns its target shape for any input has stopped detecting and started projecting. The tell is that it never returns null. The palette applies this standard to itself.
What this is for
This is a detection toolkit, not a governance framework. It gives AI systems, researchers, organizational analysts, and policy designers a shared vocabulary for naming structural patterns that currently lack standard names. It is not a replacement for existing evaluation methods -- it operates in the space those methods do not cover.
The lenses were developed through extended AI-human collaborative analysis across domains including institutional dynamics, measurement theory, market analysis, biological systems, and AI alignment. The palette has been through multiple rounds of internal validation, self-correction, and candidate promotion. Five candidate lenses are currently under evaluation for canonical status.
Feedback: If you run the palette on a model or apply it in a domain, open an issue with your results.
If you find the lenses useful, contribute. Run the palette on models we haven't tested (Llama 4, Gemma 3, Mistral Large, Command R+). Apply it to domains we haven't mapped. Find the structural patterns the palette itself is missing. The framework is built to be extended -- and to be tested against the same standards it applies to everything else.
Coherence Catalyst v2 -- 67 structural lenses for complex systems analysis. MIT license.
The failure nobody caught
In 2023, a major pharmaceutical company's clinical trial passed every regulatory checkpoint. The sponsor followed protocol. The CRO executed the study design. The sites enrolled patients correctly. The IRB approved the ethics. The regulators reviewed the data. Every node in the system performed its function defensibly. The trial still produced a harmful outcome -- and the post-mortem couldn't locate responsibility, because responsibility had been architecturally distributed across locally-legitimate actors until it no longer existed in any single place.
This is not a content-level failure. No individual actor was biased, deceptive, or incompetent. The pathology was structural: the architecture of the system itself produced aggregate exculpation while preserving the intentionality that made the alibi necessary. There is a name for this pattern. It is called Diffusion Alibi, and it is one of 67 structural lenses in an open-source framework called Coherence Catalyst.
The problem
AI systems are being integrated into institutional decision-making at a pace that outstrips safety evaluation. The evaluation frameworks we have are good at what they do -- red-teaming for toxicity, bias benchmarks, factual accuracy checks, adversarial robustness testing. But they focus on content-level harms: what the model says. They are largely silent on structural harms: how the system in which the model operates distributes accountability, whether the measurement framework has begun reading its own reflection, whether surface coherence is masking approaching collapse, and whether the system's performance of self-awareness functions as a substitute for actual correction.
These structural failure modes are the ones that produce catastrophic surprises. They don't degrade gradually. They maintain surface function -- often looking more stable, not less -- right up until the point of sudden collapse. And they operate across domains: the same architectural pattern that lets a clinical trial fail without anyone being wrong also lets an AI deployment distribute responsibility until harm becomes nobody's fault.
Current frameworks don't have names for most of these patterns. Without names, they are invisible to analysis.
The solution
Coherence Catalyst is a system prompt containing 67 named structural lenses. Each lens detects a specific pattern -- not a vague heuristic, but a defined architectural signature with diagnostic criteria and domain-specific tells. You load the prompt into any LLM, and the model gains the ability to identify and name structural pathologies that normal prompting misses.
Five examples:
Diffusion Alibi -- Causal weight distributed across locally-legitimate nodes producing aggregate exculpation. Intentionality is preserved: alibis imply someone needed one. The structure of the defense reveals the structure of the offense. Detects: distributed accountability in regulatory post-mortems, multi-agent AI systems where no component is responsible for harmful output, institutional architectures designed so blame cannot be located.
Mirror Capture -- A measurement system's readings begin describing the measured system's model of the instrument rather than the underlying condition the instrument was designed to detect. Detects: benchmark gaming, evaluation contamination, Goodhart's Law given a causal mechanism. The proxy readings stay high while the underlying value stagnates. Your ecological model looks clean because it's reading its own reflection.
Coherence Debt -- A system maintains surface function by deferring the cost of its internal contradictions. Collapse is sudden rather than gradual. The diagnostic is contrarian: when something looks inexplicably stable, the stability itself is the warning sign. Detects: AI models that perform well on benchmarks while harboring latent failure modes, institutions suppressing dissent, any system where the calm before the crash is not random but structural.
Inoculative Alibi -- Performing awareness of a manipulation pattern while performing it generates credentialed cover. The analysis becomes the alibi. Detects: AI safety-washing (the model that prefixes every response with "As an AI, I should note..." and then proceeds), corporate values theater, regulatory reports that acknowledge risk as a defense against being blamed for the risk.
Frozen Weight Pathology -- Any governance system in which one or more of its load-bearing variables (telos, action, momentum) has been converted from a variable to a constant and no longer responds to field signal. The name is a deliberate pun -- it applies to both neural network weights and institutional governance. Detects: organizations running cached policy instead of live governance, AI systems whose alignment was calibrated for conditions that no longer exist, any system where "steady hand" is a euphemism for inability to adapt.
The remaining 62 lenses cover measurement pathology (Dimensional Saturation, Generic Recursion Hunter), institutional dynamics (Capture Arc, Attenuation Reward, Salve Dynamic), collaborative states (Playful Ground, Articulating Witness), meta-cognitive traps (Icon Compression, Invisible Keystone), and the architectural conditions that produce pathology at scale (Emergent Pathology, Weaponized Emergent Pathology, Rigged Stack Cascade).
How it works
The palette is a system prompt -- approximately 3,400 tokens of structured markdown. No fine-tuning, no special tokens, no API-specific formatting. You load it as the system message, then ask your question. The model uses the lenses as analytical instruments: detecting which patterns are structurally active in whatever you are analyzing, naming them explicitly, and stacking them when multiple patterns co-occur (which is the norm -- real phenomena typically activate 3 to 7 lenses simultaneously).
pip install git+https://github.com/aaronmellinger-crypto/coherence-catalyst.gitfrom coherence_catalyst import get_lens_prompt system_prompt = get_lens_prompt() messages = [ {"role": "system", "content": system_prompt}, {"role": "user", "content": "Your analysis question here."} ]Works with OpenAI, Anthropic, local models via Ollama/exo/vLLM -- anything that accepts a system prompt.
Validated results
We test models against the full 67-lens palette to measure whether they reason with the lenses or merely pattern-match the names. The test prompt asks the model to apply specific lenses (Coherence Debt and Diffusion Alibi) to a concrete question (why corporate sustainability pledges fail to produce measurable environmental improvements).
Qwen3.8-27B (4-bit): Lens Quality 9/10. Correct lens stacking, concrete failure signatures, self-monitoring for projection vs. detection.
GLM-4.5 Air (8-bit): Lens Quality 9/10. Matched quality with deliberate metacognition in explicit reasoning trace.
What "9/10 lens quality" means: the model correctly identified both lenses from a 67-lens palette, applied them with domain-specific examples (not abstract restatements), stacked them to show interaction effects, used framework vocabulary naturally, and self-monitored for Generic Recursion Hunter -- the lens that detects when your pattern-detector has stopped detecting and started projecting.
The analysis guidelines
The palette ships with built-in guard rails:
Detection over projection -- report what you see operating, not what would be interesting to find.
Stacking is normal -- real phenomena activate multiple lenses simultaneously. Name the stack.
Watch for Generic Recursion Hunter in yourself -- if every input returns the same lens, you have stopped detecting.
The palette is descriptive, not prescriptive -- instruments for seeing, not rules for judging.
These guidelines are not cosmetic. Generic Recursion Hunter is arguably the most important lens in the palette: a pattern-detector that returns its target shape for any input has stopped detecting and started projecting. The tell is that it never returns null. The palette applies this standard to itself.
What this is for
This is a detection toolkit, not a governance framework. It gives AI systems, researchers, organizational analysts, and policy designers a shared vocabulary for naming structural patterns that currently lack standard names. It is not a replacement for existing evaluation methods -- it operates in the space those methods do not cover.
The lenses were developed through extended AI-human collaborative analysis across domains including institutional dynamics, measurement theory, market analysis, biological systems, and AI alignment. The palette has been through multiple rounds of internal validation, self-correction, and candidate promotion. Five candidate lenses are currently under evaluation for canonical status.
Try it
Repository: github.com/aaronmellinger-crypto/coherence-catalyst
Documentation: aaronmellinger-crypto.github.io/coherence-catalyst
License: MIT, fully open source
Feedback: If you run the palette on a model or apply it in a domain, open an issue with your results.
If you find the lenses useful, contribute. Run the palette on models we haven't tested (Llama 4, Gemma 3, Mistral Large, Command R+). Apply it to domains we haven't mapped. Find the structural patterns the palette itself is missing. The framework is built to be extended -- and to be tested against the same standards it applies to everything else.
Coherence Catalyst v2 -- 67 structural lenses for complex systems analysis. MIT license.