Should I quit my job?
Facts:
1) I currently work as a freelance video games marketer and I really dislike working for my main client @ $2500 USD per month. I'm on a sinking ship (the studio is failing) and morale is low; I dread logging in for work in the mornings.
2) I've just been awarded a generous $5000 USD grant from BlueDot Impact to build an AI Governance Megagame which I think has a lot of potential for the future. The grant is to develop and playtest the game for the next few months but does not include a stipend or living expenses. That said I think I can generate revenue from these playtests and in 2027 it's possible I could generate a lot of revenue from facilitating my game for institutions/events/conferences.
3) I am passionate about game design, event management, and AI governance/policy. Much moreso than I am passionate about indie games marketing/social media management.
4) I only have enough CAD cash savings to survive until ~Q1 2027. After that I am toast and I have to either become a bartender again or return home with my parents, both of which I view as fail states.
Does anyone have any advice???
UPDATE: Wow, they dropped me a day after I posted this. I guess the relationship was deteriorating faster than I realised. Would have been nice if the people in charge of me had any communication skills. Oh well. I've got other prospects. I don't plan to revert to any fail states. Thank you all for your commentary.
> I currently work as a freelance video games marketer and I really dislike working for my main client @ $2500 USD per month
That doesn't seem like much money. You could probably get a more pleasant job with better pay elsewhere, especially if you're willing to leave the games industry.
Personally given your savings situation I'd be wary of quitting to work on the game unless you have some income from it lined up in advance (eg pre-bookings for facilitation).
It's a crappy job. I mean I'm grateful because that $2500/mo saved my ass back in November, but I hate it when a contract client treats me like an employee. I won't rant about it here.
Totally willing to leave the industry. I guess part of it is a confidence issue/not believing I am experienced or knowledgeable enough to advance and get paid more for something somewhere else. But again this company is vaporising around me so I will need to get over that neuroticism pretty quickly.
As someone who also lives in Toronto and was in a somewhat similar scenario a while ago - I would say let's get you some more funding to quit! (either in booking/payment in advance for facilitation as ShardPheonix suggested, or a secondary grant source).
I feel the same way. Problems:
1) This grant already is such a crazy amount of money and I feel like I need to make more concrete progress (i.e. run the playtest) before I should court a stipend. I have vague plans in my roadmap (late Oct) for a larger Manifund proposal once I have written/photo evidence that my game actually exists, and that grant would include a modest cost-of-living stipend for a few months. Though I am worried that the project would be too commercial for Manifund's tastes; after all my optimal goal for this project is business-shaped and I get the sense they don't like investing in for-profit operations. Maybe Manifund is the wrong place.
2) I go back and forth on the payment in advance for facilitation thing. I am going to facilitate my own playtests this fall, and they'll be modestly ticketed (15 to 25 bucks a head; the only reason it's not free is because I want all 63 playres to actually show up). But after that, my hope is to stop putting together my own events and facilitate the game for organisations or at conferences. But I can't take people's money until I've tested the game 1-2 times and I feel confident it's a value creator.
Not trying to be difficult or reject your advice, haha. It's just that I've thought about this quite a lot. Certainly I think there are options/resources for me, even here in Toronto. This week I plan to throw my name in to join the Trajectory Labs coworking space for 3-6 months (I think what I do aligns with what they do), and that would meaningfully make things easier for me. Alas, I can't sleep in the office.
Anywho. I'll figure something out. Thanks for the comment.
This grant already is such a crazy amount of money
I'm not in this space, but a single $5000 grant seems pretty small if you're expected to work on this full time.
True. Particularly because that $5000 grant isn't for my income, it's entirely for project expenses like printing components and renting venues.
Maybe that was hyperbolic of me. It's just a little exciting because I've never undertaken a solo project and received the funding to make it happen before, so even this small grant feels like a lot. It's like a big hot potato that fell into my hands and if I drop it I face a clawback lawsuit and/or a severe loss of trust/confidence from the AI governance community.
Oh yeah, I agree that it's a big deal that someone thought your project was important enough to provide that kind of funding! I just also think if the expectation is for you to work on this full time, it's entirely reasonable to ask for even more.
Well, you'd definitely see me if you come into Trajectory :-) But i think the manifund proposal is probably a very good idea. Also, dw, there are plenty of commercial flavoured things on Manifund - Austin of Manifund is also known to care about profitability as a metric even for non-profit success.
Also see Austin on finding funding fast.
1) I currently work as a freelance video games marketer and I really dislike working for my main client @ $2500 USD per month. I'm on a sinking ship (the studio is failing) and morale is low; I dread logging in for work in the mornings.
Yes.
Congrats on the grant! That does sound like an exciting opportunity. As for the rest of your situation, sorry to hear that your job is dragging you down. Being stuck in such a situation sucks, and I guess the way forward depends on many details that us commenter don't have much context about. Like, would returning to your parents be worse than staying in this job for longer?
Quitting with only a few months of runway sounds pretty risky to me, and chances are that much of the time that you would want to focus on working on your game, you'd be stressed out about your financial situation and where you'll be at two months later (at least I would).
Maybe some other questions one could raise:
From what you've described, I'm not overly optimistic about any of these, but who knows.
Overall, I think something along the lines of "hang in there as long as necessary while at the same time looking for alternatives (whether a more enjoyable and/or better paid job, or more funding for your project)" would seem wise to me.
I'm not sure how much you weigh enjoyment of work compared to having money to live, but my take is that quitting your job is to go all in on the game for 2 months is worth it.
- Best case: You're game takes off, you're working on things that you're passionate about and you generate a lot of money from it. Authentic passion seeps into work and how you talk about your work, which influences how much work you get in these types of industries. The bluedot grant is a good signal of your game imo, and if the stakes are high one could assume you will give it your all (so increased likelihood of success). You have to go hardcore though, in addition to making an great game you may need to aggressively market/network irl and online.
-Worst case: You spend two months working on something cool and impcactful instead of something that gives you dread, whilst being roughly net 0 loss? 2 months based on a rough assumption of you losing $2500 x 2 from the client, in favor of the $5000 grant. Then after that you can figure out what to do, and Bluedot could point you in other directions? Even if you enter one of your two fail states, it is possible it'll fire you up to escape them as opposed to continue waking up with dread in your current job after not having tried at all.
I can't imagine being passionate about event management, but the people who are get snapped up by the Generator Residency https://generatorresidency.org/
It's a weird interest. I think it's the same reason I'm so interested in film production and media publishing and the hospitaliy industry and video game development. I guess my brain just enjoys interdepartmental logistics. I don't think I'd be a very good wedding planner, though.
I actually threw my name into Generator last time, didn't get picked up. To be fair, though, I think it was a weak application. I was basically just showing them that one blogpost. I've made a lot more progress since then and I have serious goals I want to hit by Christmas. So maybe another time if I can get US authorisation.
I've found LLMs useful as an assistant for coding, writing, research, bizdev advice, and general ideation, but even the almighty Fable 5 is currently very unhelpful to me as an amateur game designer.
I might write a whole blog post about this. For some reason LLMs just don't seem to understand what makes a game mechanic "good"? When I present my existing ruleset to Fable 5 and ask it to create new content or design a new mechanic according to specifications, it never has good ideas.
In the set of disciplines in which I have some working knowledge (which is admittedly very few), game design has by far the most significant taste gap between skilled humans and the best LLMs.
But again, I'm just an amateur. I'm curious what professional game designers would say to this.
I'm not a professional game designer but I've found Claude useful for a hobby game project, though it's probably in a very different genre than yours - a vaguely Long Live the Queen/Magical Diary-ish visual novel/time management sim. I've asked it stuff like
The last one is probably closest to yours. While it didn't give me an amazing idea for a replacement right away, I found that discussing different possibilities with it eventually helped me figure out my own idea.
Interesting question - do AIs have a feel for fun?
AIs are good at recommending existing fun products: movies, games and trips I’d like, based on what I tell it about my taste. But that’s mechanical pattern matching and memorization.
AIs can also figure out what makes things un-fun: nuisances, unfairness, confusion, alienation, and propose practical fixes.
But when I‘m designing a novel experience (event planning, a software product), the intuition for the experiential kernel that would make it fun, meaningful, satisfying, consistently comes from me.
"coding, writing, research, bizdev advice, and general ideation"
Depending what you mean by general ideation, these all basically require memory recall and low-IQ recombination of existing memories and techniques. The actual intelligence required is extremely low (though would maybe be much higher for a human who doesn't natively reason the way LLMs do, the same way calculators can beat us at multiplication problems).
But if you ask AI to do any task that requires real creativity or intelligence, it flops. Game dev design, novel programming problems that require judgement, writing jokes, writing stories anybody would want to read, writing scripts for social media content anyone would want to watch, coming up with app ideas, doing research, etc. These things all require Actual Intelligence, i.e. the ability to efficiently navigate concept space by using judgement and forming new judgements to refine the search, and LLMs don't have much of it. Maybe their "true" IQ is 10, to the extent IQ makes sense for an AI model. But they fake it with memory the same way humans fake walking ability based on evolutionary memory.
Joke writing is my favorite personal benchmark, because the results are somewhat objective for a personal benchmark. You laugh at the AI output or you don't. I have never had an AI model yet that could write funny jokes at all except by accident. Fable was the first model where I was often able to detect a faint hint of something. Some of the premises felt like they had once been in the same room as an actual joke. But the model is still unfunny. The first funny model will probably be the one that kills us, because it implies intelligence that also unlocks all that other stuff, so I'm paying attention to this.
I'm working on a 2d grid-based roguelike and this is especially true in my case, in part because AIs are still bad at vision and spatial reasoning and therefore can't fully understand what they're making. My game was basically unplayable for Fable until it made itself a text-based harness, and even then it was quite bad at doing basic navigation.
YMMV! Not sure why you're being downvoted.
It's possible that you are better at AI prompting than me.
It's possible that my current game design is quite complex, to the extent that LLMs have a harder time writing content or designing mechanics for my game as opposed to, say, prompting it to "write an expansion for Monopoly".
What's your experience with LLMs and game design?
wrote a sibling :)
i suspect we are having similar outcomes here, but that i'm noticing "hey, this is better than the last one: its ideas are sometimes interesting to consider" while you're noticing more "ehh, it still doesn't get it".
certainly my best results come from isolating a specific problem, and asking the model to summarize the problem back to me. unclear whether this is more useful than a diary! well, except in the critical sense that i am not motivated to use a diary.
i find that past a fairly low amount of complexity, the model starts forgetting rules, or ignoring/inventing context.
(as for the downvotes: i mean, if you're sitting at +X agreement, i should be at -X, no?)
to be clear, my bar for 'impressed' is pretty low, having tried previous models for these things.
i got good results from fable by asking the model to come up with a system that described a few mechanics, and then look for missing mechanics within that system. i was pleased with the system it developed, and found a few of its proposals inspiring, though not usable verbatim.
I've also found that LLMs are terrible at understanding existing game mechanics and giving suggestions on builds or decks, even when the information is clearly spelled out on the internet. My guess is that relatively few gaming rules are included in the training data because it doesn't make the models better at code?
So I've been on LW for about a month. For the record, yes, I have read the new user's guide because I'm a slavish rule-follower. I've also browsed the very extensive concepts index.
I am still not entirely clear on what exactly does one post on LessWrong?
I get that there's a distinction between Front Page posts and Personal Blog posts. But at the same time, whether you make a FP/Personal post, my intuition is that there are certain types of posts you wouldn't post on LW.
But sometimes that intuition is challenged when I see very non-ratty posts like this gross mac and cheese recipe. (Unless maybe the recipe is supposed to be epistemologically profound in its radical simplicity?)
Surely there must be some kinds of posts that would be taken down or at least downvoted into oblivion because they're a bad fit. Thinkpieces about B2B sales? Educational pieces about population ecology? Personal posts about the quirks of raising a toddler? Kirby Super Star: Ultra fanfiction?
Posts that get downvoted into oblivion usually have problems like AI writing, trying too hard to make persuasive arguments, mischaracterizing things, or generally doing things stylistically that annoy people.
If people feel like your post is off-topic (with no other problems) it will usually just get ignored, not downvoted.
It's really hard to say what on-topic actually is here though, and you could probably get away with any of those posts if they're written in the right way.
There are no hard rules, but I see it as a sequence of filters.
First, like Brendan said, there are some mistakes that result in negative karma. If you can avoid this, feel free to post.
After you pass the first filter, there are various things that could bring positive karma. It can be research, or something personal, or fiction... You only need to succeed in one of those criteria. If you fail at all of them, karma will be close to zero.
Finally, moderators decide whether to move your post on front page, which I guess is a combination of topic and karma.
Sometimes the karma depends on what other people write, for example if we start having too much fiction, people will be annoyed even by things they would have normally liked or ignored.
My guess is that the population-ecology and parenting advice posts would score well controlling for quality and author recognition, but the B2B sales advice and Kirby fanfic would get downvoted unless there was some kind of thematic local twist. Only one way to find out!
"It's over, Haltmann," said Kirby. "You've lost. Disable the Access Ark."
Haltmann had lost a lot of blood but could still move and speak. "No," he croaked.
"Do it or I'll kill you," said Kirby.
Haltmann laughed. "A threat? You insult my knowledge of decision theory, Kirby. The Access Ark only responds to my biosignature. If you want me to do something for you, you'll have to offer me something I want."
Kirby shrugged and ate Haltmann. "Great, now that I have his biosignature—"
A turret on the wall shot balls of some sort of sticky substance at Kirby. It smelled like peanut butter. "What—what—" Kirby gasped, as his throat began to swell up.
"And his peanut allergy?" said a cold yet feminine metallic voice.
I just attended a practice run of D. Scott Phoenix's "The Endgame" milsim/wargame/LARP event in Berkeley, CA. It was interesting. If you take our simulation tonight as gospel, xAI is going to win the AI compute/fab race IFF China blockades Taiwan.
~40 people split into teams representing different actors in the current AI space (Anthropic, OpenAI, xAI, but also Venture Capital, the US, China, and The Public, etc.) and engaged with each other over rounds, brokering deals with each other and taking actions to see what happens in a simulation of the near future.
China and the US went to a hybrid war, the US nationalised AI labs and American silicon fabs, Anthropic and Deepmind merged, xAI become the foremost silicon provider in the country, and some other stuff happened.
Oh yeah, and Claude GigaMythos leaked, political travel plans went public, and a few US politicians got assassinated. But in the end, the US and Chinese governments and their over-reliance on AI for governance and policy led to an AI-enforced world peace!
It was a cool experience, if a little flawed. So much potential. Unfortunately it seems like Scott is going to present this game this week at a conference and then never run it again.
I think if you gave me 14 days and a few playtests I could design a far better version of this game that flowed better, was more engaging, and achieved the goal better (helping powerful people in AI model what the future might look like based on incentives). I understand the need to keep the game tight and lightweight, but the game really needs better modelling for mass media, social media, relationships, resources, and "what can I do on a given turn?". There are also some changes I would make to the structure and the factions available.
(I'm not critiquing out of disrespect. It was just my hope that Scott might hire me to work on and improve his game. Alas, it looks like The Endgame concludes in the next few days.)
I might write a LW effortpost about this actually. I love game design and I would one day like someone to pay me to design games/run events for them. I am particularly inspired by other simulation games like the UK based Megagame Makers who have been doing this kind of thing for years.
Thank you Scott for the invite! :D

Insider info from the Inkhaven Writer's Residency @ Lighthaven: we're being given swanky enamel Prestige Pins with the Inkhaven logo on them.
But not just anyone gets a Prestige Pin.
The pins were created to encourage us to spread our creative wings and try different things. In order to earn a pin, you must have published an Inkhaven post that falls into EACH of seven categories: Fiction, Emperical, Informational, Persuasive, Humour, Advice, and Personal.
The program was announced a couple of days ago. I had already written in 6/7 posts and I was planning to write a humour-ish post anyways, so last night I requested a pin and this morning the team signed off on it. I am now the owner of an Inkhaven Prestige Pin.
Woo!
This makes me think, though. Was this an experiment, and have I been had?
I think part of the ethos of the Inkhaven Residency¹ is to cultivate and encourage agency and self-motivation through a high pressure environment. Having published 32,958 words at Inkhaven at the time of writing, I'm pretty astonished with how much I've output here. I never would have done this if I wasn't shoved into a weird compound in Berkeley with 54 53 other weirdos.
I am quite pleased with this. I feel like I had more of a "you can just do things" drive a few years ago when I was in university, and I really think this month might be getting me out of that creative rut.
But if the purpose of Inkhaven is (in part) to curate our internal locus of agency, isn't it odd that we were suddenly given a very clearly external motivator— a enamel flame for us moths to fly to? It feels a bit odd to me.
After all, many of us came to Inkhaven with very clear writing objectives: some people here write exclusively about AI safety, AI policy, or AI technical stuff. Some of us are travel and lifestyle bloggers, and some of us are etymologists. Curation of one's voice is another motif of the experience, and by forcing ourselves to write fiction or sature when that's very much not our niche... is that productive?
Maybe. Certainly many talents and niches aren't fully realised until they are thrust upon us. The enamel pin is a target for us archers to hit, but the real learning happens as we take aim.
But at the same time it's a little funny that this high stakes program of self actualisation introduced the Prestige Pin two thirds of the way through. What an inversion of our personal creative expression to hand us a bullet point list and give us a "while supplies last!" marketing pitch.
It certainly worked on me.²
On Saturday the organisers are holding an open exhibition for the Residency— the Inkhaven Fair— here at Lighthaven. Come check it out. You might get to watch all the Prestige Pin awardees be publicly humiliated for being sheep rather than Real Bloggers.
[1] - I have spent weeks wondering what exactly the ethos of this program is— a program on which Lightcone Infrastructure reportedly loses tens of thousands of dollars on each time they run it. I have thoughts on this but I won't fully write it up until May.
[2] - I'd again like it on record that I had already met most of the requirements for the pin. But maybe my only "humour" post wouldn't have been written in quite the same style if there wasn't a Prestige Pin on the line...
If Lesswrong had certain prestige pins, like Pokemon badges, I would write on LW more often. The idea of a trinket that expresses social status to a niche group of people in-the-know, is so sticky to my brain.
If LW did this, I think they could sell the badges for $30-50 each, and you only 'unlock' those purchases in the store when you - for example - have 10 posts hit the front page, or have made 100 useful wiki edits, or got to 1000 Karma.
Oh, and then they could do special badges for Petrov Day participants! And April fools! And my god, a Shoggoth Enamel pin would be one of my most prized possessions (Only available to purchase during Fooming Shoggoth concerts).
I'm going to go research enamel pin creation now.
P.S. Can we see a photo of your pin?!
Be in awe of my Prestige.

I do like where your head's at, though. As a new LessWrong user who loves nothing more than in-group signalling, I'm a little sad at the total lack of LW/Lightcone merch available.
I bow at your majesty. It is a lovely pin, sir.
I have a Redbubble store, where I sell some Rationalist merch. All the Rationalist stuff (expect for a few items I was too lazy to change the setting on) are sold at cost, and I make no money from them. However, they're honestly not very good, and I mainly put the rationalist stuff up there, because I wanted to a "notice confusion" phone case.

I think it's ~silly that Lightcone doesn't have a merch store, because I'd buy the shit out of Lesswrong merch. Like, I also imagine Shoggoth Blind boxes, like PopMart. Which I would easily spend $200 on to buy all an entire carton of (if it's guaranteed that the carton contains all the variants)

I think "Lightcone Brand Paperclips" is also another cute, untapped market.

Shoggoth Pin Mockup: And according to Custom Ink it's only $909 USD to get 150 of these shipped to me in Australia! If I was rich, I would do this.

Why do we treat writing as a particular sacred cow, even among us LLM users? Plenty of us (myself included) embrace LLMs for any number of applications, but we (myself included) often spit on AI generated writing and get annoyed when people try to pass it off as their own, explicitly or implicitly. It really bugs me when someone replies to an email of mine with transparently Claude-generated emails, or when blog posts have AI smell.
I don't really get the epistemics of this. Why writing, which is arguably the most basic function of an LLM?
My main guesses are:
1) There's no way to present a piece of writing as a distinct encapsulated artifact. You can human-write a post and say "hey, here's some code I had Claude generate" but a blogpost has no top-level context framing, if that makes sense?
2) It feels like it's some violation of trust in our relationship if you give me a 1000 word post or document to read, but you generated it in 20 seconds with a prompt.
There might be better reasons, or better phrasings of my reasons. But I just find it strange that in 2026, where Fable 5 actually writes pretty decent prose a lot of the time, us in-the-know AI users tend to frown on LLM-smelling text.
EDIT: I should clarify that I am very happy that I have this reflexive allergy to LLM text. I just don't know why I feel that way.
I don't get the downvotes, I think it's an important question to ask when I arises.
To me, finding undisclosed LLM text just is a very strong indicator the person publishing it wants to get away with sloppy work and doesn't respect (the time and epistemic environment of) their readers very much. Clearly this doesn't apply to every case, it's just by far the most common occurrence, so by posting LLM text, you're moving into a different reference class. The most likely scenario behind someone posting LLM text, purely on very general priors, is that they wrote a short prompt and got a slop response. If the person invested more effort into the interaction and has reason to believe their LLM text is of high quality and worth reading, then they should point out why (which involves being transparent about LLM usage in the first place).
I think I'd still have this preference if there was only a single LLM in existence which was both aligned and superintelligent. Text is supposed to have effects on readers, so being transparent about the nature of a text just seems like the right move.
To me, finding undisclosed LLM text just is a very strong indicator the person publishing it wants to get away with sloppy work and doesn't respect (the time and epistemic environment of) their readers very much.
Yeah, that is exactly how I perceive it.
Even if you let the LLM generate the answer, the least you can do is tell it "be concise" or "give me a summary" or quote the relevant part instead of dumping the entire wall of text. If you don't bother doing even that, that signals you clearly do not value my time.
Another thing: prompts matter. If you don't share the prompt, I have no idea whether you asked your AI to "compare X and Y, evaluate pros and cons from various perspectives, and then make a conclusion" or "give me 7 reasons why X is better than Y". The former is potentially helpful, the lattes is not. And the more AI-psychotic the person seems, the more likely I think it was the latter.
Or doing philosophy, but AI-generated philosophy is harder to evaluate and could be confounded by the cultural hegemon's priors, RLHF favoring sycophantic philosophy, etc.
If I ask you a question, and you have your friend answer but pass it off as your own answer, I will lose trust in you. That's deception even if you agree with your friend's wording.
If I hire you to do a job, and you subcontract it out without telling me, that's breach of contract.
Writing isn't special, writing is first. This will all happen with robots too. If I hire you to fix something, and a robot shows up, that's going to be a big problem.
I think the subcontractor/robot situation really depends. There are cases where I hire someone purely so that something gets done, without caring about who does it. If they send another person, or a robot, I may be entirely fine with that. But there are cases where this doesn't apply, say when I hire a person because they specifically have a great reputation or I otherwise think highly of their work and I care a lot about details and quality.
I think everyone in the developed world should endeavour to eat more plants and fewer animals. There are many motivators to this, and I think the two strongest motivators are greenhouse gas emissions and animal suffering. I think the vegetarians and vegans are probably broadly morally correct, and I think that any rationale that motivates one to vegetarianism or veganism is a valid one.
That said, I'm extremely skeptical of any attempt to create a metric to somehow create a proportional relationship between animal suffering hours and tons of C emitted. That seems extremely difficult to me, even when it's mapped out in neat tables.
I am a layman with no education or training on AI or AI Safety. I have been following the Claude Mythos/Glasswing arc as well as the general explosive improvement in AI coding ability. I believe LLMs are still exploding with competence, particularly in coding.
I think the reasons for coding being the #1 area of LLM performance improvement are:
1) There's a clearer "correct answer" for training a bot to write a quicksort implementation than, say, a short story
2) There's more enterprise money in training coding bots than, say, short story bots
All that said, I find it hard to believe that AGI, if and when it comes, will be LLM-shaped?
As a user of LLMs for 3.5 years (and GPT sandbox before ChatGPT came out), I feel as though there are certain areas— writing and rhetoric in particular— where the models are approaching some sort of ceiling. I'm not impressed with Claude 4.7 in that capacity, nor was I with 4.6. They feel about as good as 4.5. I can say the same of ChatGPT, Grok, and to some extent Gemini.
But this is just vibes! And maybe a small amount of motivated reasoning as a self-proclaimed blogger-fictionist. We don't really have a good way to index writing ability, especially not fiction. All we have is the opinions of people with taste, and the taste-makers still seem to say AI-written fiction is crap. I tend to agree.
What I'm trying to get at is that when people talk about Mythos as AGI, I'm like, "yeah, it's a super smart coder and it's going to change the world of software and cybersecurity forever, but also it's just an LLM." Maybe the AGI of the future will have an LLM component, but I can't help but cringe when people say whatever new LLM could be AGI.
I don't know. Again, I am a layman. But of all the people who follow AI and LLMs, I'm probably in the bottom quartile of the ranking of people who think about LLMs from a software dev standpoint.
My take about the writing quality stagnating while hard verifiable metrics go up is the delineation between reinforcement learning and increasing the size of the base model. Unfortunately the specific details of model training are private so it makes claims like this feel a bit hollow, but it seems likely that eg Opus 4.6 is "the same model" as Opus 4.5 just with many more steps of RL applied to make it better at specific discrete verifiable tasks. This way of thinking predicts that notable scaleups in the size of the base model, which is presumably what Mythos is, would have notably better skills in the soft/unmeasurable areas such as writing, contextual understanding, humor etc. This is supported by the anecdotal reports at the end of the Mythos model card, though it's hard to truly say without having the model ofc. It's certainly POSSIBLE that LLMs will never be better at writing fiction and related soft fields than they are today, but I doubt this and think the wall you see is a result of the above RL trend rather than a true ceiling in capabilities being reached anytime soon.
I guess my main gripe with people who have your attitude re "LLMs can never be AGI" is like, what concretely is missing to you? What are the differences between Mythos and what to you would unequivocally be AGI, and why do you believe you need some mysterious secondary component in order to acquire that capability? To me the most obvious answer is memory and the ability to actively learn ("continual learning"), but one could even argue that is unneeded since the main character in Memento is certainly a general intelligence is he not?
I've been trying to think about substances lately less as vices/fun things to do/habits, and more as tools.
Caffeine is a tool. As is alcohol or THC or CBD. There are other tools out there such as psilocybin or LSD or GHB that I haven't tried yet but could also be useful to me.
The binary of "recreational" versus "non-recreational" drugs just isn't a super useful binary to me.
Take Benadryl (diphenhydramine) for example. I don't get allergies but I sometimes take 100mg diphenhydramine when my sleep schedule is fucked up and I need to sedate myself.
But just like any tool, substances can be abused. Diphenhydramine can be abused with short-term and long-term side effects, as can acetaminophen, or aspirin, or indeed THC or alcohol or caffeine.
This is a pretty big mindset shift for me. I'd like to be more intentional and conscientious in this phase of my life, and being critical of my morning coffee/evening beer rituals is probably smart.
If substances are like tools, then addictive drugs are like that kind of software you can't uninstall.
Glosso is very new platform, so much remains to be seen, but it's so far failed to pass the Metaposting Filter: in order for a platform to have any viability as a new social network, the proportion of posts talking about the platform (metaposts) has to be a very small %.
Bluesky barely escaped this fate (and is still limping on). Vidme succumbed to this filter.
Obviously it's not causal— there are many reasons why Vidme failed— but this is a phenomenon that I've noticed.
Metaposts are a symptom that a platform is in a state of flux. People talk about things that are novel or controversial, and when your platform is novel or controversial (and therefore it's what people are talking about), that's a sign of a platform in dire straits (or at least one that has yet to settle down and find its niche in the information ecosystem).
To be clear, not all the public posts on Glosso are metaposts. But majority in my feed are. I will believe more strongly in Glosso as a platform when the public page is populated mostly with posts of actual substance.
As a normie-adjacent, it seems to me like we are 0-3 years away from some 9/11-style AI-assisted disaster that wakes up the people of the world. Mass casualties, enormous damage, economic crisis, and/or something else. Something that doesn't wipe out humanity, but is extremely alarming and universally understandable.
I'm frightened by what this catastrophe might be, but it seems to me not only the most probable scenario, but also the best scenario. The alternative is frogboiling.
The Hugging Face thing was crazy and is rocking my world, but basically nobody in my life, my family, or my community knows anything about it. I think it's just too abstract to explain to someone cold.
But a severe, headline-friendly crisis affecting the western world seems highly likely to me. That's probably the first and best chance for the great powers to cooperate and create a global AI treaty. The actual thought-through, researched, and wargamed policies need to be ready when that time comes.
I'm not a shape rotator. I don't have any role to play in technical alignment. But now that I'm without income for the time being and I have some limited resources to work with, I'm going to be shifting my focus to governance, policy, and communication for a little while by making my side project my main project. Maybe it'll have some positive effect somewhere. I want to publish a larger post on this subject shortly.
It's interesting to me just how many people in the AI research/safety crowd still think AI is worse at writing than the best humans. I think it's true, but it's an interesting contrast to the AI bros on Twitter who are like "I just made this in 10 minutes with [TOOL]. [INDUSTRY] is dead."
Naive proposal as a nontechnical non-researcher. I'm assuming either that this has been done or there's a reason why it's infeasible/not worth doing.
Let's create some kind of complex virtual world, a detailed (but limited) model of the human world.
Let's populate it with many fairly dumb, earlier gen AI agents who represent humans. They think they are people. Most of them have normal lives and jobs but a small subset of them are AI researchers working on frontier models. Let's then take a much smarter AI and throw it into the environment as if the agents just created what is, to them, ASI.
For GPT-3 agents, a Claude Fable 5 instance would appear to be an ASI.
Then we just see what happens.
It'd be like Google's "Smallville" except the starting moment is the instant the ASI instantiates itself inside of the Smallville AI Training Lab (or whatever) and starts learning about its world.
Has anyone tried that? Does it work/produce meaningful information? I've been reading about Smallville last night/this morning but I'd appreciate any further reading.
I think the eval awareness point is the big problem like you say (it would be very hard to convince a frontier AI that this world is real), but I don't think it's necessary for the AI under test to believe that it's not an AI?
Has anyone here ever played the Intelligence Rising game made by Shahar et al? If so, can I ask you about your experience?
Increasingly it feels like wearing swimwear is actually more obscene than just swimming naked.
Like I'm not a "nudist" but lately when I go for a dip in a pool or lake or spa, and I'm wearing swim trunks, I'm just thinking "what is even the point of this?" It kind of just worsens the swimming experience. It makes it more of a hassle, it feels uncomfortable, and, yeah, to the prudes I'm tempted to argue that a bikini is more obscene than skinny dipping.
The purpose of a bathing suit is to cover "private parts" while still being practical enough to swim in (to include drying off afterwards). Fine. But as fashion and practicality trend towards more and more revealing swimwear (mostly for women), the only remaining fabric serves only to "cover up" at the absolute limits of what's socially acceptable. At that point, what purpose does the swimwear still serve? To me, a small bikini is just actually highlighting what it's trying to cover up. It's emphasising the body, except for the parts that the wearer/the culture considers to be shameful.
A nude body has none of that. Nothing is highlighted or faux-hidden. It just is.
I'm not saying that I have a particular interest in making swimming "less obscene", FWIW. My interest is mostly practical. I visited a lagoon/spa in Iceland and while it was lovely, I felt myself thinking "man, the fact that I have these swim trunks clinging to me is probably the worst sensory part of this entire experience."
I really don't understand the argument against nude beaches in my native North America. They're quite rare here, and "normal people" don't go to them. "Nudists" go to "nude beaches". That's not the case in many other parts of the world, but alas, NA is weird about it.
Why are some people moralised by nude beaches but not by bikinis/speedos?
Why are some people moralised by nude beaches but not by bikinis/speedos?
I think they are/were? Speedos were considered gay in USA a few decades ago, based on a reasoning that if a guy wants to be on the beach wearing anything less covering than a spacesuit, the only logical explanation is that he wants to expose his body to other guys. Or something like that, not sure about the details (am European).
Even a Type III civilisation that has solved AI alignment and FTL travel faces Great Filter-adjacent problems; these problems might have consequences anywhere from capping further intergalactic expansion to existentially destroying the civilisation.
Very very interesting Glosso question that I'd like to pose to LessWrong:
Do you think there is at least a 1% chance that the claims of the Bible (particularly of Jesus in the Gospels and the Acts of the Apostles) are true?
Why or why not?
Bringing the dead back to life? I don't think so. (Though people can be mistaken about whether someone's dead.)
But the miracle of the loaves & fishes is just a church potluck, or Stone Soup for that matter. Of course you end up with more leftovers than the amount of food you started with — once you've got that kind of party going, people keep showing up and bringing more.
I'm not sure how to interpret the probability here, but if I was presented with a slot machine showing a random historical instance of someone claiming a miracle occured with a +1% payout if the miracle was actually magic, I'd pull the lever all day. Religious books tend to make very little sense, we haven't discovered any evidence of magic, and we have plenty of evidence that people can have religious experiences without magic (or can just lie).