Traditions are sets of practices that have survived and passed down through time from generation to generation. They usually weren't arrived at through a process of reasoning, and even if they were, that reasoning is often lost to time. And yet, the very fact that these traditions propagated through time gives us information about their utility... (read more)
AI Control in the context of AI Alignment is a category of plans that aim to ensure safety and benefit from AI systems, even if they are goal-directed and are actively trying to subvert your control measures. From The case for ensuring that powerful AIs are controlled:.. (read more)
The Open Agency Architecture ("OAA") is an AI alignment proposal by (among others) @davidad and @Eric Drexler. .. (read more)
Singluar learning theory is a theory that applies algebraic geometry to statistical learning theory, developed by Sumio Watanabe. Reference textbooks are "the grey book", Algebraic Geometry and Statistical Learning Theory, and "the green book", Mathematical Theory of Bayesian Statistics.
| User | Post Title | Wikitag | Pow | When | Vote |
Handoff is when a group of humans grant AIs a position of trust, decision-making, or both.
A paradigm example would be a frontier AI company handing off to their own AIs. These AIs might be in charge of: prioritising between research agendas; deciding how to develop and deploy successor AIs; managing relations with governments (foreign and domestic); managing relations with clients, suppliers, competitors, and the public; managing philanthropic ventures; etc.
Handoff varies along several dimensions:...
What a nice explainer this has been so far! I do think the proof given for Bayes' rule in odds form has room for improvement. Not sure if it shows up in every 'path', so I'll link the specific url I visited here: https://www.lesswrong.com/w/bayes-rule-odds-form?pathId=61b&lens=introduction-to-bayes-rule-odds-form
(sadly I seem to be unable to propose edits directly)
The proof involved deriving three equalities:
1.
P(Hj)/P(Hk) * P(e0|Hj)/P(e0|Hk) = P(e0∧HJ)/P(e0∧Hk)
so far so good. This derives directly from the definition of conditional probability, P(Y)*P(X|Y)=P(X∧Y), when applied to both the numerator and denominator of the fraction on the left hand side.
2.
P(e0∧Hj)/P(e0∧Hk) = (P(Hj|e0)/P(e0)) / (P(Hk|e0)/P(e0))
This looks like a roundabout application of the definition of conditional probability. While it technically holds, it seems to me like it follows less directly from P(X∧Y)=P(Y)*P(X|Y) than the following would:
P(e0∧Hj)/P(e0∧Hk) = (P(Hj|e0)*P(e0)) / (P(Hk|e0) * P(e0))
3.
the final equality would then change from
(P(Hj|e0)/P(e0)) / (P(Hk|e0)/P(e0)) = P(Hj|e0) / P(Hk|e0)
to
(P(Hj|e0)*P(e0)) / (P(Hk|e0) * P(e0)) = P(Hj|e0) / P(Hk|e0)
Putting it all together into a single line:
P(Hj)/P(Hk) * P(e0|Hj)/P(e0|Hk) = P(e0∧HJ)/P(e0∧Hk) = (P(Hj|e0)*P(e0)) / (P(Hk|e0) * P(e0)) = P(Hj|e0) / P(Hk|e0)
Typo: "causally downstream of the moment the moment"
Untrusted monitoring is the AI control technique where the untrusted AI's actions are sent to some AI-based monitoring process which is itself untrusted. Usually, the monitor is another instance of the same AI, then if one of them is scheming then they both are. This contrasts with trusted monitoring, where we have a specific reason to believe its reports will always be honest (e.g. if it's a particularly weak AI whose alignment we can directly certify).
Untrusted monitoring has the advantage that the strong monitor can keep up with the convoluted actions the untrusted AI might be taking, as its capabilities scale well past any trusted model. The central failure mode is collusion, where the monitor intentionally reports bad actions as "good". Research in untrusted monitoring seeks to:
Interesting. Another aspect or hidden layer would be intonation, e.g.
Level 1: “There’s a LION across the river.” = Exclusive emphasis on the Lion as the direct object/entity.
Level 2: “There’s a lion ACROSS the river.” = Expressing direction/orientation
Level 3: “There’s a lion across THE river.” = Intonation emphasis on declarative article as the river being exclusive (vs There's a Lion across A river" where the direct object is an indeterminate river)
Level 4: “There’s a lion across the RIVER.” = Exclusive emphasis on river as the direct object.
Consider an action with benefits but also large costs. If a proponent of the action doesn't feel conflicted and grapple with the costs, that's evidence that they don't appreciate the costs and thus don't have an informed view on the action. (And even if the action is basically correct, the version of the action they endorse likely isn't the optimal version of the action;version; they probably won't do a good job of minimizing the costs.) The failure to feel conflicted is called a "missing mood." This term comes from Bryan Caplan's The Missing Moods.
Real-world experience provides an excellent foundation for continued professional success. level 6 nvq construction offers a practical way to demonstrate management capability through evidence collected from active construction projects. The qualification values competence developed through hands-on responsibility and leadership. This workplace-focused approach continues to benefit construction professionals throughout their careers. The College of Contract Management also provides construction-related programmes that encourage lifelong learning.SPAM
Handoff is when a group of humans grant AIs a position of trust, decision-making, or both.
A paradigm example would be a frontier AI company handing off to their own AIs. These AIs might be in charge of: prioritising between research agendas; deciding how to develop and deploy successor AIs; managing relations with governments (foreign and domestic); managing relations with clients, suppliers, competitors, and the public; managing philanthropic ventures; etc.
Handoff varies along several dimensions:
The central questions include:
See also:
Consider an action with benefits but also large costs. If a proponent of the action doesn't feel conflicted and grapple with the costs, that's evidence that they don't appreciate the costs and thus don't have an informed view on the action. (And even if the action is basically correct, the version of the action they endorse likely isn't the optimal version; they probably won't do a good job of minimizing the costs.) The failure to feel conflictedabsent attitude is called a "missing mood." This term comesmissing mood," from Bryan Caplan's The Missing Moods. There can also be missing moods other than not appreciating costs, and the same concept applies to incorrect moods that are present — Caplan's examples include pacifists sympathizing with evil regimes and libertarians sneering at the poor.
The most commonUntrusted monitoring is the AI control technique is monitoring: all where the untrusted AI's actions are sent to some AI-based monitoring process before they are escalated to humans. By default, thiswhich is itself untrusted. Usually, the monitor is itself untrusted, unlessanother instance of the same AI, then if one of them is scheming then they both are. This contrasts with trusted monitoring, where we have a specific reason to believe its reports will always be honest (e.g. if it's a particularly weak AI whose alignment we can directly certify).
For example, the monitor might just be a another instance of the same AI. ThisUntrusted monitoring has the advantage that itthe strong monitor can keep up with the convoluted actions itthe untrusted AI might be taking, as its capabilities scale well past any trusted model. The central failure mode is collusion, where itthe monitor intentionally reports bad actions as "good". UntrustedResearch in untrusted monitoring research seeks to:
In all countries of the world slavery is outlawed at least this much (with tolerance of penal slavery being common) but in practice illegal slavery happens anyway. The Global Slavery Index (wikipedia link) has information on this.
Consider an action with benefits but also large costs. If a proponent of the action doesn't feel conflicted and grapple with the costs, that's evidence that they don't appreciate the costs and thus don't have an informed view on whether the action is worthwhile. The failure to feel conflicted is called a "missing mood." This term comes from Bryan Caplan's The Missing Moods.
Consider an action with benefits but also large costs. If a proponent of the action doesn't feel conflicted and grapple with the costs, that's evidence that they don't appreciate the costs and thus don't have an informed view on the action. (And even if the action is basically correct, the version of the action they endorse likely isn't the optimal version of the action; they probably won't do a good job of minimizing the costs.) The failure to feel conflicted is called a "missing mood." This term comes from Bryan Caplan's The Missing Moods.
Consider an action with benefits but also large costs. If a proponent of the action doesn't feel conflicted and grapple with the costs, that's evidence that they don't appreciate the costs and thus don't have an informed view on whether the action is worthwhile.action. The failure to feel conflicted is called a "missing mood." This term comes from Bryan Caplan's The Missing Moods.
Untrusted monitoring is the AI control technique all where the untrusted AI's actions are sent to some AI-based monitoring process which is itself untrusted. Usually, the monitor is another instance of the same AI, then if one of them is scheming then they both are. This contrasts with trusted monitoring, where we have a specific reason to believe its reports will always be honest (e.g. if it's a particularly weak AI whose alignment we can directly certify).
The most common AI control technique is monitoring: all the untrusted AI's actions are sent to some AI-based monitoring process before they are escalated to humans. By default, this monitor is itself untrusted, unless we have a specific reason to believe its reports will always be honest (e.g. if it's a particularly weak AI whose alignment we can directly certify).
For example, the monitor might just be a another instance of the same AI. This has the advantage that it can keep up with the convoluted actions it might be taking, as its capabilities scale well past any trusted model. The central failure mode is collusion, where it intentionally reports bad actions as "good". Untrusted monitoring research seeks to:
If a Son of CDT agent goes on to create further agents, all of those agents will have the same magic moment. They will all care about whether or not Omega's knowledge of them is causally downstream of the moment the moment the CDT agent first wrote Son-of-CDT code.
Why does it show as deleting all the LaTeX in my edit? i only added (Assume the veterinarian is correct.) The LaTeX does not seem to be deleted on the actual page.