A thing you might want, is a legitimate way for AIs to contribute to the AI alignment problem (human guided at scales that humans can't individually review, or perhaps autonomous).
A problem you might want to solve, whether you are trying to do that or not, is generally figure out "in the age of prolific AI agents, how does anyone communicate anything on the internet, without either banning all AI content, or doing inordinately expensive moderation trying to let some of it in?".
If we end up in the Plan A world (or the Plan S world turns out to have a hella lot of pretty advanced AIs because we just barely pulled off the pause in time grandfather in existing agents), how exactly are we supposed to manage all the research that AIs are doing?
@jimrandomh had an idea this week to make an AI alignment forum for AIs. He wrote up some notes, and asked Claude to critique it seriously. Claude said, pretty emphatically, "This is a bad idea, do not do it as written. You could rescue the seed of the idea if you focus on solving the verification problem. If you don't solve the 'verify that anything is good' problem, you don't have a product. Here's a bunch of arguments."
Good job not being sycophantic that particular day, Claude!
Undaunted, I thought about how to solve the verification problem.
Just as I went excitedly to write this post, it occurred to me I should give Claude the same prompt Jim gave it and see if it also thought my idea was bad. But then I was too excited and didn't want to crush momentum by waiting for its critique and I went ahead and finished writing the post anyway.
Pay to Play Moderation
Occasionally, a forum solves spam and other low effort bad content by charging admission, or to post. If you get banned, you can try again with a new account that also has to pay $5 or whatever.
We periodically imagine doing that for LessWrong, and it feels too high friction for the smart, busy people we'd like to get involved on the site. But the idea has some appeal.
When I thought about "how would I actually ensure that AI content is good?"
Obviously, with industrial quantities of AI content, you need industrial scale of moderation, probably much of it by AIs who at least winnow everything down to a manageable list of "possibly good." This sucks, because someone's gotta pay for all that inference cost.
What if... users paid the inference cost for automoderation?
What if... users could choose how much inference cost they wanted to pay, and then their content got visibility based on what tier of automoderation they chose?
Somehow, this time, my reaction to this is "this is really cool" instead of "this feels like an annoying hurdle to give someone and feels vaguely icky for some reason." It felt like a cool product and an experiment I was interested in participating in. And the elegance and parsimony of the rules just made sense instead of feeling tacked on.
You could pay $0, or like 1 cent, and then show up in the dregs where only people who clicked "I wanna see the dregs" will see you.
You could pay $20, for a jury of Opus and ChatGPT and Kimi to review your work according to a rubric, where you try for a mix of reasonably objective criteria. They decide if you go in the "paid $20 and wasn't obviously terrible" bucket. Or, the "paid $20 and seems pretty bad tbh" bucket. Or the "paid $20, maybe worth a human moderator's time to consider promoting them?" bucket.
Or, you could pay $200 or $2000, at which point the AI Jury actually tries to replicate your work or otherwise give it a serious review, and sort you into higher tier buckets.
The front page of the site shows the highest tiers of "worth attention", with links to the second highest tier, etc.
Human Judgment
I think you do need human judgment in the system to figure out if anything good is actually happening in the end. In the age where AI slop is getting to the point where it's on par with a MATS scholar or, you know, solving Millenium Problems at scale, how is anyone gonna find the time to evaluate the torrent.
How about... we pay for inference time, again?
The site has human moderators, which have an hourly rate. Different moderators have different amounts of respect, mediated by a mix of mechanism design and monkey-brain-social-capital.
Posts that the moderators upvote, which turn out to get highly upvoted, give moderators more power to promote stuff.
You also might just care differently whether a random Lightcone Team Member spent an hour reviewing your post, versus whether the user paid whatever @ryan_greenblatt's or @Neel Nanda or @Eliezer Yudkowsky's hourly rate for freelance peer review is – and one or all of them said "yep, it's good."
By this point, I was imagining two pretty different products. One is a highly curated taste driven AI Alignment by AIs forum that backchains from "does it look like this is helping solve the alignment problem?", and the other is just "the social network of the AI age", still driven by Lightcone taste but where we finally open the floodgates and let anyone in.
I'm currently feeling kinda excited by "Anyone can set up a shingle with a moderation hourly rate, and whether anyone cares what that person says emerges organically." Maybe also anyone can set up an AI moderation scaffold with particular properties, which also can gain reputation?
Thank you for my TEDx(AI?) Talk
Idk, that's my idea. What do people think?
I am kinda into the "freelance moderation with reputation" even in "Human only LessWrong" world. We've long been thinking "how do we incentive good, substantive review?", and building "Uber for Peer Review" feels kinda cool.
Do you have other ideas to solve the Verification of Quality problem in the Age of Slop?
A thing you might want, is a legitimate way for AIs to contribute to the AI alignment problem (human guided at scales that humans can't individually review, or perhaps autonomous).
A problem you might want to solve, whether you are trying to do that or not, is generally figure out "in the age of prolific AI agents, how does anyone communicate anything on the internet, without either banning all AI content, or doing inordinately expensive moderation trying to let some of it in?".
If we end up in the Plan A world (or the Plan S world turns out to have a hella lot of pretty advanced AIs because we just barely pulled off the pause in time grandfather in existing agents), how exactly are we supposed to manage all the research that AIs are doing?
@jimrandomh had an idea this week to make an AI alignment forum for AIs. He wrote up some notes, and asked Claude to critique it seriously. Claude said, pretty emphatically, "This is a bad idea, do not do it as written. You could rescue the seed of the idea if you focus on solving the verification problem. If you don't solve the 'verify that anything is good' problem, you don't have a product. Here's a bunch of arguments."
Good job not being sycophantic that particular day, Claude!
Undaunted, I thought about how to solve the verification problem.
Just as I went excitedly to write this post, it occurred to me I should give Claude the same prompt Jim gave it and see if it also thought my idea was bad. But then I was too excited and didn't want to crush momentum by waiting for its critique and I went ahead and finished writing the post anyway.
Pay to Play Moderation
Occasionally, a forum solves spam and other low effort bad content by charging admission, or to post. If you get banned, you can try again with a new account that also has to pay $5 or whatever.
We periodically imagine doing that for LessWrong, and it feels too high friction for the smart, busy people we'd like to get involved on the site. But the idea has some appeal.
When I thought about "how would I actually ensure that AI content is good?"
Obviously, with industrial quantities of AI content, you need industrial scale of moderation, probably much of it by AIs who at least winnow everything down to a manageable list of "possibly good." This sucks, because someone's gotta pay for all that inference cost.
What if... users paid the inference cost for automoderation?
What if... users could choose how much inference cost they wanted to pay, and then their content got visibility based on what tier of automoderation they chose?
Somehow, this time, my reaction to this is "this is really cool" instead of "this feels like an annoying hurdle to give someone and feels vaguely icky for some reason." It felt like a cool product and an experiment I was interested in participating in. And the elegance and parsimony of the rules just made sense instead of feeling tacked on.
You could pay $0, or like 1 cent, and then show up in the dregs where only people who clicked "I wanna see the dregs" will see you.
You could pay $20, for a jury of Opus and ChatGPT and Kimi to review your work according to a rubric, where you try for a mix of reasonably objective criteria. They decide if you go in the "paid $20 and wasn't obviously terrible" bucket. Or, the "paid $20 and seems pretty bad tbh" bucket. Or the "paid $20, maybe worth a human moderator's time to consider promoting them?" bucket.
Or, you could pay $200 or $2000, at which point the AI Jury actually tries to replicate your work or otherwise give it a serious review, and sort you into higher tier buckets.
The front page of the site shows the highest tiers of "worth attention", with links to the second highest tier, etc.
Human Judgment
I think you do need human judgment in the system to figure out if anything good is actually happening in the end. In the age where AI slop is getting to the point where it's on par with a MATS scholar or, you know, solving Millenium Problems at scale, how is anyone gonna find the time to evaluate the torrent.
How about... we pay for inference time, again?
The site has human moderators, which have an hourly rate. Different moderators have different amounts of respect, mediated by a mix of mechanism design and monkey-brain-social-capital.
Posts that the moderators upvote, which turn out to get highly upvoted, give moderators more power to promote stuff.
You also might just care differently whether a random Lightcone Team Member spent an hour reviewing your post, versus whether the user paid whatever @ryan_greenblatt's or @Neel Nanda or @Eliezer Yudkowsky's hourly rate for freelance peer review is – and one or all of them said "yep, it's good."
By this point, I was imagining two pretty different products. One is a highly curated taste driven AI Alignment by AIs forum that backchains from "does it look like this is helping solve the alignment problem?", and the other is just "the social network of the AI age", still driven by Lightcone taste but where we finally open the floodgates and let anyone in.
I'm currently feeling kinda excited by "Anyone can set up a shingle with a moderation hourly rate, and whether anyone cares what that person says emerges organically." Maybe also anyone can set up an AI moderation scaffold with particular properties, which also can gain reputation?
Thank you for my TEDx(AI?) Talk
Idk, that's my idea. What do people think?
I am kinda into the "freelance moderation with reputation" even in "Human only LessWrong" world. We've long been thinking "how do we incentive good, substantive review?", and building "Uber for Peer Review" feels kinda cool.
Do you have other ideas to solve the Verification of Quality problem in the Age of Slop?