I have spent some time over the last few months engaging with moderately informed undergraduates on AI risk (mostly by teaching AI risk arguments in my courses along with enough background material to give them some grounding). If others are interested, I would happily share my takeaways at greater length but I wanted to share a couple of interesting things:
1) Today's college students do not think AI progress has been rapid. Generally, GPT-4 maxed out their internal benchmarks (they don't know it by name obviously but that's the referent), chatbots haven't gotten much better since then as they see it, and 2023 was eons ago in college student time. Working through a project with Claude Code changes nothing because they mostly have zero intuition for what that would have looked like in the past. Their naive mental model if they project progress forward is incremental improvement.
2) My students tend to think that catastrophic risk arguments are designed as (or at least function as) distractions from the "real" issues (jobs, environmental impact, etc.) and that they should not be taken seriously as a result. Nearly all of my students would be very happy to pause AI today, but none of them in response to catastrophic risk.
3) When forced to grapple with the arguments seriously, different discussions with my students independently converged on the view that existential risk was probably a good thing because ordinary people otherwise have zero leverage over AI. That is, if AI leaders genuinely believe there is existential risk (and that, therefore, they will personally die), then this might be a way to get them to stop developing job-killing AI that will otherwise take down the rest of us.
If others are interested, I would happily share my takeaways at greater length
Please do! (I find your students' perspective alien and would be interested to read more.)
I'm currently an undergraduate student and have tried to spark discussions related to long-term AI risks in e.g. computer science Discord servers for my university, I've gotten much of the same results. For example, I shared AI 2040 without comment and got responses like:
"Glad people are trying to think about the future of generative AI, and they have clearly put a lot of effort into this, but I'm always skeptical of these long term plans. They are also making the assumption that generative AI will exponentially scale, which - while possible, is just pure speculation and blustering. I mean seriously, "Wildy Superintelligent - Qualitatively beyond human science and comprehension"? As the technology currently stands, they aren't intelligent in any way, they are prediction machines. And like with anyone talking about GenAI, these people are likely just trying to sell it. Most of the people at the top have a vested interest in GenAI and are trying to sell it as a tool that will outperform all of humanity and lead us to a prosperous utopia if controlled correctly"
I offered the obvious pushbacks and we went back and forth but ultimately only got to the common ground of the regulations that are clearly warranted now even given the more immediate risks (i.e. treating it like other high-risk dual-use technologies with negative externalities). In general I imagine this will be the most effective strategy for making actual incremental political progress. You can lead the horse to the reasoning behind rapid progress in AI capabilities forecasts but you can't make it drink.
What subset of undergraduates exactly are you engaging with? My own college experience is that the set of undergraduates who would cheerfully get into politics-adjacent arguments with the professors are heavily tilted towards certain views.
I’d be curious to hear what your students think about the possibility that AI leaders don’t take existential risk seriously, but that everybody dies anyway. Is that outcome a good thing?
Working through a project with Claude Code changes nothing because they mostly have zero intuition for what that would have looked like in the past.
This is a problem I've never thought about but it makes total sense! e.g. I can't perceive how big a difference GPS/Google Maps has made because I never used paper maps.
This just seems false to me though? I never knew a time without telephones, but it is intuitively obvious to me that their invention was a big deal. Like, I can just try to imagine a telephone-less world. In my head. Humans can do that.
To look at an llm-less world I don't even have to do that, I can just visit my grandparents, who've probably never used an llm in their lives.
I think the point is more that it's very difficult to estimate the difficulty of something youve never done and especially to differentiate between things you might be able to do yourself in, say, a few hours vs things requiring serious effort or expertise.
Most magic tricks, for example, are pretty simple conceptually and it's easy to think "oh if I spent the weekend on it, I could that." Sometimes that's true but other moves that look very similar from the outside take hundreds of hours of practice to pull off.
This is sort of how I was expecting The Public to react to the release of ChatGPT, and then it didn't, and instead actually appeared to be paying some attention and at least sort of understand that this was a big deal. So I went and updated my model of The Public on this to be more optimistic. And now The Public is reacting to the agents like this, and I feel blindsided in the opposite direction. I guess my problems with predicting The Public go pretty deep and can't be fixed just by adjusting some global first moment variable.
Interesting that they don't feel the progress. Maybe schoolwork shaped things were mostly saturated years ago, at least within their abilities to discriminate?
" and 2023 was eons ago in college student time."
It may be that, in objective terms (years) they predict progress at the same rate you do. Its just that what they consider a "long time" is different.
Slow, incremental progress to a person who considers 3 years to feel like "eons", might look like lightning progress to an adult who expects their life in 3 years to look basically the same as it does now.
Imagine a homeless person - they would worry about food resources or accommodation rather than longer term risks. Similarly, if someone still needs to worry about survival before long term risks come, they would probably pay more attention to that. Unfortunately it is true that there are still many near term things that could affect shorter term survival to a lot of people, so even though x risks are something to be aware of/paying attention to, it likely won’t be the only thing to worry about. This situation is usually a product of the era and environment, and seems more prevalent in places with greater economic/power/social disparity.
On a side note, I have been reflecting recently and felt that power dynamics/abuse, and bad actors are probably the common central factors in both longer term risks like x risks and other current risks.
I 100% agree with 1: I think the only solution I have to this is more centralised workshops where students are forced to only use good old Stack Overflow and do things by hand; in many industries this is still required.
2 + 3 are super interesting; I've taught undergrad level chemistry and physics about a year ago and my perception was a bit different - students were very quick to spot quite worrying hallucinations (I guess since the domains were less representative in the training set of normal chatbots), to the point where some were just very reluctant to use them at all for anything, including coding.