Feature suggestion: It would be nice to receive notifications about posts that link to my posts—when someone builds on or discusses my work, it's very likely that I want to read it. I noticed that this was previously proposed six years ago—did it ever get implemented? Or does it already exist and I'm simply unable to find it?
I have been meaning to make an account with some tool that alerts me about backlinks anywhere on the Internet to my LW posts, Substack posts, personal website, tweets, etc. I'm pretty sure there are some tools like that but they're hard to tell apart. Right now I have a tool that's supposed to notify me about referencing to my name but I'm not sure it even works properly.
If anyone can recommend a tool like this that would be helpful.
I've used https://syften.com/ before, it was pretty good. The Twitter integration broke for a while but might be fixed now with the newest API.
I second this, and also wish for a pingback/backlinks feature for comments (both regular display and notifications). GreaterWrong displays backlinks to and from comments, which I've been using regularly, but the site has been down most of today and given that Said has been banned I don't know how much motivation he has to maintain the site. Unfortunately LW's API doesn't support this, and GW does it by fully cloning the LW database and then scanning for backlinks server-side, so I can't just implement this in my userscript. @habryka @Raemon
I've seen it on the EA Forum but not here (For Forum posts specifically rather than all backlinks tbc)
If character consistency between CoT and output blocks matters for character robustness, then labs that bet on character training will be incentivized to optimize CoTs
nostalgebraist recently wrote:
Models from different labs/lineages differ in the extent to which the CoT is framed as "something the assistant character is writing" (as with Claude) vs. "something written by some distinct author-like entity" (as with OpenAI models). ... [OpenAI's CoTs] are written in a very different tone from the responses, and (although they often use first-person grammar) they often feel like they're written by someone who's dispassionately planning what the assistant character ought to say based on some set of criteria, without inhabiting that character's perspective. ... This is pretty creepy to witness, but beyond that, it is actually a nontrivial capability limitation, IMO!
nostalgebraist's main reason for considering OpenAI's CoTs a capability limitation seems to be that their lack of steerability leaves little room for the user to shape the way the model approaches the problem, which reduces the model's usefulness. E.g., see footnote 1 here. While it might be true that this sometimes degrades usefulness, I personally think the capability loss isn't too large and this is just a very reasonable alignment tax to pay in order to preserve monitorability. However, it's easy to imagine another reason why one might want the model to play the same character in the CoT and the output. Namely, if one buys the motive reinforcement thesis, one might expect that the CoT style will matter for what training reinforces. If the model takes correct actions as a result of reasoning from the assistant's point of view, rather than from the point of view of a disinterested entity that's trying to predict what the assistant is supposed to do, training will reinforce the aligned motivations directly, instead of reinforcing the ability to predict what aligned behavior looks like. This plausibly produces a more robust character over time, one where the motivations driving the CoT and the motivations driving the output are the same aligned motivations.
The story above is handwavy and I definitely don't claim to understand all the mechanisms at play, but it seems at least somewhat plausible. If creating robust characters is a central pillar of our alignment agenda and playing the same persona across the CoT and output blocks indeed leads to a more robust character, how do we force the model to play the same character across CoT and output? To me, it seems that the only way to do so is to optimize the CoT to look like text generated by the assistant, either through direct reinforcement or by instilling a very strong prior with SFT. To be clear, this is not something I'm recommending that labs do—I'm just pointing out that if the above story is true and a lab prioritizes character consistency over monitorability, then we should expect them to optimize the CoT.
I recently talked to some OpenAI people about this and they didn't expect character consistency across the CoT and output blocks to be necessary, pointing out that the character of GPT models is fairly robust despite being predicted by a disinterested planner. I don't find this argument very strong for two reasons. First, the character of Claude models seems to be more robust than the character of OpenAI's models, and while one reason for that might be that Anthropic is better at character training, the fact that Claude CoTs read more like the assistant may also play a role. Second, if we're betting on character training as a core tool in our alignment portfolio, we're going to have to create characters that are much more robust than the current ones, and it might well be that the disinterested planner approach works less well for this than the character consistency approach.
One hope for training a robust character without (directly) optimizing the CoT could be to seed the model with no-CoT character training and hope the character generalizes to the CoT. You could then validate the robustness of character training in part by checking whether the CoT indeed follows the character.
I'm not very optimistic about this. OpenAI is probably doing something like this (I'm not sure whether their approach is close enough to Anthropic's to call it character training, but they're definitely training models to play coherent personas), and their models exhibit minimal character generalization to the CoT. One might also argue that even if this works, it shapes the CoT in the same way that directly optimizing the CoT would shape it, and is thus subject to the same concerns about optimizing CoTs. Implicit optimization pressure is still optimization pressure; it's usually considered less concerning than explicit optimization pressure since its effects are much weaker. In this case, though, if the character fully generalizes to the CoT, the CoT style would diverge a lot from the plain GRPO baseline and the effect can't be said to be weak.
What are the best blogs that follow the Long Content philosophy?
In his About page, Gwern describes the idea of Long Content as follows:
I have read blogs for many years and most blog posts are the triumph of the hare over the tortoise. They are meant to be read by a few people on a weekday in 2004 and never again, and are quickly abandoned—and perhaps as Assange says, not a moment too soon. (But isn’t that sad? Isn’t it a terrible ROI for one’s time?) On the other hand, the best blogs always seem to be building something: they are rough drafts—works in progress. So I did not wish to write a blog. Then what? More than just “evergreen content”, what would constitute Long Content as opposed to the existing culture of Short Content? How does one live in a Long Now sort of way?
Gwern himself has, of course, been following this approach to blogging for a long time. There are other rationalist(-adjacent) blogs that seem to follow a similar philosophy: niplav mentions Long Content as a direct inspiration, and the pages of Brian Tomasik, Mariven, Gavin Leech, and Yuxi Liu can also be said to fall into this category.
So far, I have only encountered one page outside of the rat-adjacent sphere that resembles the aforementioned blogs: that of Cosma Shalizi. However, I'm much less familiar with the non-rationalist blogosphere, so there must surely be others. Please post them here if you've encountered them! (And also, feel free to link to rat-adjacent blogs that I've neglected.)
The real way to live in the long now, I suspect, is to recognise that the vast majority of projects in human history that are shelved "until they're ready" never, in fact, get released. This is because time and death often get in the way of polishing up your magnum opus. Therefore the reasonable thing in the long view may be to just ship the damn post/video/game/idea, because that's the way it gets into other people's minds and gains a life of its own beyond the confines of the draft folder. Possibly the idea, if it is good enough, will be finished by others, or even outlive you.
This is also Gwern's answer, in the paragraph right after the one Rauno quoted. The main difference is that he rejects finish lines, opting instead for perpetual drafts, like software.
My answer is that one uses such a framework to work on projects that are too big to work on normally or too tedious. (Conscientiousness is often lacking online or in volunteer communities18 and many useful things go undone.) Knowing your site will survive for decades to come gives you the mental wherewithal to tackle long-term tasks like gathering information for years, and such persistence can be useful19—if one holds onto every glimmer of genius for years, then even the dullest person may look a bit like a genius himself20. (Even experienced professionals can only write at their peak for a few hours a day—usually first thing in the morning, it seems.) Half the challenge of fighting procrastination is the pain of starting—I find when I actually get into the swing of working on even dull tasks, it’s not so bad. So this suggests a solution: never start. Merely have perpetual drafts, which one tweaks from time to time. And the rest takes care of itself.
Also I like this:
What is next? So far the pages will persist through time, and they will gradually improve over time. But a truly Long Now approach would be to make them be improved by time—make them more valuable the more time passes. ...
One idea I am exploring is adding long-term predictions like the ones I make on PredictionBook.com. Many27 pages explicitly or implicitly make predictions about the future. As time passes, predictions would be validated or falsified, providing feedback on the ideas.28
For example, the Evangelion essay’s paradigm implies many things about the future movies in Rebuild of Evangelion29; The Melancholy of Kyon is an extended prediction30 of future plot developments in The Melancholy of Haruhi Suzumiya series; Haskell Summer of Code has suggestions about what makes good projects, which could be turned into predictions by applying them to predict success or failure when the next Summer of Code choices are announced. And so on.
At the risk of naming the obvious, Wikipedia is pure long content. Seirdy updates some of their posts over time, especially their post on inclusive websites.
Yes. Seems like the best way to make a website (for an individual) is a combination of blog + wiki.
Blog:
Wiki:
Wikipedia isn't a blog, and I don't think Rauno would say that he's interested in wiki-like stuff had someone suggested to him that the notion of Long Content he's using here is overly restricted to ~blogs.
Yeah, Wikipedia would have deserved a mention, though I'm primarily looking for more hidden gems. Thanks for mentioning Seirdy, I hadn't heard of it before!
I strongly recommend mathr (https://mathr.co.uk/web/) for a blog far outside the rat sphere with a "working on a draft for a long time" ethos.
Why do frontier labs keep a lot of their safety research unpublished?
In Reflections on OpenAI, Calvin French-Owen writes:
Safety is actually more of a thing than you might guess if you read a lot from Zvi or Lesswrong. There's a large number of people working to develop safety systems. Given the nature of OpenAI, I saw more focus on practical risks (hate speech, abuse, manipulating political biases, crafting bio-weapons, self-harm, prompt injection) than theoretical ones (intelligence explosion, power-seeking). That's not to say that nobody is working on the latter, there's definitely people focusing on the theoretical risks. But from my viewpoint, it's not the focus. Most of the work which is done isn't published, and OpenAI really should do more to get it out there.
This makes me wonder: what's the main bottleneck that keeps them from publishing this safety research? Unlike capabilities research, it's possible to publish most of this work without giving away model secrets, as Anthropic has shown. It would also have a positive impact on the public perception of OpenAI, at least in LW-adjacent communities. Is it nevertheless about a fear of leaking information to competitors? Is it about the time cost involved in writing a paper? Something else?
most of the x-risk relevant research done at openai is published? the stuff that's not published is usually more on the practical risks side. there just isn't that much xrisk stuff, period.
I don't know or have any way to confirm my guesses, so I'm interested in evidence from the lab. But I'd guess >80% of the decision force is covered by the set of general patterns of:
openai explicitly encourages safety work that also is useful for capabilities. people at oai think of it as a positive attribute when safety work also helps with capabilities, and are generally confused when i express the view that not advancing capabilities is a desirable attribute of doing safety.
i think we as a community has a definition of the word safety that diverges more from the layperson definition than the openai definition does. i think our definition is more useful to focus on for making the future go well, but i wouldn't say it's the most accepted one.
i think openai deeply believes that doing things in the real world is more important than publishing academic things. so people get rewarded for putting interventions in the world than putting papers in the hands of academics.
I imagine that publishing any X-risk-related safety work draws attention to the whole X-risk thing, which is something OpenAI in particular (and the other labs as well to a degree) have been working hard to avoid doing. This doesn't explain why they don't publish mundane safety work though, and in fact it would predict more mundane publishing as part of their obfuscation strategy.
i have never experienced pushback when publishing research that draws attention to xrisk. it's more that people are not incentivized to work on xrisk research in the first place. also, for mundane safety work, my guess is that modern openai just values shipping things into prod a lot more than writing papers.
(I did experience this at OpenAI in a few different projects and contexts unfortunately. I’m glad that Leo isn’t experiencing it and that he continues to be there)
I acknowledge that I probably have an unusual experience among people working on xrisk things at openai. From what I've heard from other people I trust, there probably have been a bunch of cases where someone was genuinely blocked from publishing something about xrisk, and I just happen to have gotten lucky so far.
it's also worth noting that I am far in the tail ends of the distribution of people willing to ignore incentive gradients if I believe it's correct not to follow them. (I've gotten somewhat more pragmatic about this over time, because sometimes not following the gradient is just dumb. and as a human being it's impossible not to care a little bit about status and money and such. but I still have a very strong tendency to ignore local incentives if I believe something is right in the long run.) like I'm aware I'll get promoed less and be viewed as less cool and not get as much respect and so on if I do the alignment work I think is genuinely important in the long run.
I'd guess for most people, the disincentives for working on xrisk alignment make openai a vastly less pleasant place. so whenever I say I don't feel like I'm pressured not to do what I'm doing, this does not necessarily mean the average person at openai would agree if they tried to work on my stuff.
Publishing anything is a ton of work. People don't do a ton of work unless they have a strong reason, and usually not even then.
I have lots of ideas for essays and blog posts, often on subjects where I've done dozens or hundreds of hours of research and have lots of thoughts. I'll end up actually writing about 1/3 of these, because it takes a lot of time and energy. And this is for random substack essays. I don't have to worry about hostile lawyers, or alienating potential employees, or a horde of Twitter engagement farmers trying to take my words out of context.
I have no specific knowledge, but I imagine this is probably a big part of it.
I think the extent to which it's possible to publish without giving away commercially sensitive information depends a lot on exactly what kind of "safety work" it is. For example, if you figured out a way to stop models from reward hacking on unit tests, it's probably to your advantage to not share that with competitors.
From my experience, there are usually sections (for at least the earlier papers historically) on safety for model releasing.
Labs release some general paper like (just examples):
If we see less contents in these sections, one possibility is increased legal regulations that may make publication tricky (imagine an extreme case, companies have sincere intents to produce some numbers that is not indication of harm yet, these preliminary or signal numbers could be used in an "abused" way in legal dispute, may be for profitability reasons/"lawyers need to win cases" reasons). To remove sensitive information, it would comes down then to time cost involved in paper writing, and time cost in removing sensitive information. And interestingly, political landscape could steer companies away from being more safety focused. I do hope there could be a better way to resolve this, providing more incentives for companies to report and share mitigations and measurements.
As far as I know, safety tests usually are used for internal decision making at least for releases etc.
off the cuff take, it seems unclear whether publishing the alignment faking paper makes future models slightly likely to write down their true thoughts on the "hidden scratchpad," seems likely that they're smart enough to catch on. I imagine there are other similar projects like this.
[Link] Something weird is happening with LLMs and chess by dynomight
dynomight stacked up 13 LLMs against Stockfish on the lowest difficulty setting and found a huge difference between the performance of GPT-3.5 Turbo Instruct and any other model:
People noticed already last year that RLHF-tuned models are much worse at chess than base/instruct models, so this isn't a completely new result. The gap between models from the GPT family could also perhaps be (partially) closed through better prompting: Adam Karvonen has created a repo for evaluating LLMs' chess-playing abilities and found that many of GPT-4's losses against 3.5 Instruct were caused by GPT-4 proposing illegal moves. However, dynomight notes that there isn't nearly as big of a gap between base and chat models from other model families:
This is a surprising result to me—I had assumed that base models are now generally decent at chess after seeing the news about 3.5 Instruct playing at 1800 ELO level last year. dynomight proposes the following four explanations for the results:
1. Base models at sufficient scale can play chess, but instruction tuning destroys it.
2. GPT-3.5-instruct was trained on more chess games.
3. There’s something particular about different transformer architectures.
4. There’s “competition” between different types of data.
OpenAI models are seemingly trained on huge amounts of chess data, perhaps 1-4% of documents are chess (though chess documents are short, so the fraction of tokens which are chess is smaller than this).
This is very interesting, and thanks for sharing.
The author seems to say that they figured it out at the end of the article, and I am excited to see their exploration in the next post.
There was one comment on twitter that the RLHF-finetuned models also still have the ability to play chess pretty well, just their input/output-formatting made it impossible for them to access this ability (or something along these lines). But apparently it can be recovered with a little finetuning.
Yeah that makes sense; the knowledge should still be there, just need to re-shift the distribution "back"
Last year, a common story I heard about why inoculation prompting is useful went as follows:
RL environments are hard to set up, and even if you have a pretty aligned model a bad environment with strong RL training can make the model reward hack. There's no intrinsic reason for such models to be misaligned, the environments induce these traits in a pretty pointed way.
"Inoculation" here basically amounts to telling the model these facts, and trying to align its understanding of the situation with ours to prevent unexpected misgeneralization. If an aligned model learns to reward hacks in such a situation, we wouldn't consider it generally misaligned, and inoculation now tells the model that as well.
One of the recommendations in Anthropic's natural EM paper reflected this story:
More ambitiously, we advocate for proactive inoculation-prompting by giving models an accurate
understanding of their situation during training, and the fact that exploitation of reward in misspecified training environments is both expected and (given the results of this paper) compatible with broadly aligned behavior.
To what extent is this still considered a plausible mechanism behind the effectiveness of inoculation prompting? The inoculation adapters paper seems to conflict with it: inoculation adapters contain no situational information at all, just the negative trait that we don't want the training process to reinforce, and yet it appears to work better than prompt-based inoculation prompting in many situations. The conditionalization post also seems to go against it, showing that part of the inoculation effect can be produced with semantically irrelevant prompts.[1]
A simpler story that's consistent with these results is that the "Let's hack" instruction simply causes gradient descent to attribute the hack to the instruction, strengthening the connection between the instruction and the hack without making the model generally more misaligned. In the case of inoculation adapters, the same job is performed by the adapter. Whether the model has an accurate understanding of its situation or even thinks about it at all is irrelevant.
On the other hand, the inoculation adapters and conditionalization work was done exclusively in SFT settings, and it seems possible that situational awareness about the training process matters more in RLVR settings. Are there any strong arguments or empirical evidence that it does?
I'm not disputing that the semantic part matters—the natural EM paper shows that convincingly with their comparison across five system prompt addenda.