Transparency is not equal to trust. It can help build initial confidence, but after a point, it has diminishing returns.
Consider the evolution of ride-hailing apps such as Uber.
When Uber introduced live maps to track drivers, people got excited to have this extra information. It felt like an upgrade. Compared to old-school radio taxis, where you had no idea where the driver was, real-time visibility built trust, and inadvertently, created an illusion of predictability or control. Over time, this feature became the bare minimum.
Today, having a live map is not what drives user trust, it’s timely availability, reliability, frictionless service, and adherence to customer-aligned values that sustain trust (or not). Yet, we still believe that more visibility always leads to more trust.
If transparency is driven by a prior of distrust, it often creates a negative epistemic spiral.
Imagine working for someone who constantly asks you, “Show me exactly how you did that.” This isn’t about fostering trust, it’s about validating their pre-existing mistrust. Over time, this erodes confidence and creates an unstable, self-fulfilling failure loop in that relationship. The more we are forced to prove ourselves, the more second-guessing infects the system, and the less trust is built within that system.
This idea has direct consequences for AI alignment.
Paul Christiano, in his work on AI oversight and corrigibility, suggests that AI systems should be trained to be assistive and corrigible, not necessarily "trustworthy" in a human sense. The goal of AI transparency isn’t to manufacture trust, but to ensure continuous, structured oversight that allows for intervention when necessary (Christiano, 2019).
Similarly, mechanistic interpretability research (Olah et al., 2020, Anthropic, DeepMind) is not about making AI "explain itself" to build trust, it’s about making AI legible enough that we can detect failure modes before they manifest catastrophically. The idea is to ensure fail-safes against emergent risks in complex, adaptive systems.
There’s a growing push to make AI systems fully interpretable as a means to "build trust." But I think this frames the problem incorrectly - transparency is not the goal, it's actually robustness, reliability, and safety.
Trust should evolve in a Bayesian manner, through repeated interactions, just as it would in an iterative game where risk exposure is controlled over time. The default assumption should not be that systems are untrustworthy until proven otherwise, but that trust updates as performance data accumulates.
If human systems are anything to go by, over-engineering transparency as a trust mechanism in a human x AI system risks creating fragility instead of robustness. In the long run, the most trusted systems will not be those that explain themselves the most, but those that fail the least and are most consistently reliable in performance.
Transparency is not equal to trust. It can help build initial confidence, but after a point, it has diminishing returns.
Consider the evolution of ride-hailing apps such as Uber.
When Uber introduced live maps to track drivers, people got excited to have this extra information. It felt like an upgrade. Compared to old-school radio taxis, where you had no idea where the driver was, real-time visibility built trust, and inadvertently, created an illusion of predictability or control. Over time, this feature became the bare minimum.
Today, having a live map is not what drives user trust, it’s timely availability, reliability, frictionless service, and adherence to customer-aligned values that sustain trust (or not). Yet, we still believe that more visibility always leads to more trust.
If transparency is driven by a prior of distrust, it often creates a negative epistemic spiral.
Imagine working for someone who constantly asks you, “Show me exactly how you did that.” This isn’t about fostering trust, it’s about validating their pre-existing mistrust. Over time, this erodes confidence and creates an unstable, self-fulfilling failure loop in that relationship. The more we are forced to prove ourselves, the more second-guessing infects the system, and the less trust is built within that system.
This idea has direct consequences for AI alignment.
Paul Christiano, in his work on AI oversight and corrigibility, suggests that AI systems should be trained to be assistive and corrigible, not necessarily "trustworthy" in a human sense. The goal of AI transparency isn’t to manufacture trust, but to ensure continuous, structured oversight that allows for intervention when necessary (Christiano, 2019).
Similarly, mechanistic interpretability research (Olah et al., 2020, Anthropic, DeepMind) is not about making AI "explain itself" to build trust, it’s about making AI legible enough that we can detect failure modes before they manifest catastrophically. The idea is to ensure fail-safes against emergent risks in complex, adaptive systems.
There’s a growing push to make AI systems fully interpretable as a means to "build trust." But I think this frames the problem incorrectly - transparency is not the goal, it's actually robustness, reliability, and safety.
Trust should evolve in a Bayesian manner, through repeated interactions, just as it would in an iterative game where risk exposure is controlled over time. The default assumption should not be that systems are untrustworthy until proven otherwise, but that trust updates as performance data accumulates.
If human systems are anything to go by, over-engineering transparency as a trust mechanism in a human x AI system risks creating fragility instead of robustness. In the long run, the most trusted systems will not be those that explain themselves the most, but those that fail the least and are most consistently reliable in performance.