CNBC is reporting that Leopold Aschenbrenner's Situational Awareness fund was forced to sell its public stock holdings to Citadel after heavy losses in the market. It is unclear if the fund sold its stake in Anthropic, a private company. The article also states that Aschenbrenner is engaged to Avital Balwit, who holds a position as Dario Amodei's chief of staff, but I don't think this information is relevant in general or to this incident in particular.
From my perspective as an uninformed spectator, I would expect the bets made by Situational Awareness to be positive EV over the long run. The fund has demonstrated incredible performance until this point and I think their portfolio would have continued to perform well. I can only speculate that the fund's losses were caused by scaling their bet sizing and/or leverage too aggressively.
While I agree there's reason to think traditional software will decline in the long run and the AI infra buildout will continue, I think this is being too kind to Situational Awareness. Their leverage implies the bet was that this trend will continue nearly monotonically, with the amount of wiggle room they considered possible being miniscule, implied by the amount of leverage relative to their bankroll. Those distinctions are everything in investing - sizing is the bet! In other words, this was a bet on the path AI and traditional software equity valuations would take, not just a bet on their destination.
From Matt Levine's always excellent Money Stuff newsletter today:
This is a naturally long-term trade. Aschenbrenner’s famous June 2024 essay series is titled “Situational Awareness: The Decade Ahead.” The point is not, like, “SK Hynix will beat earnings expectations next quarter”; the point is stuff like “by the end of the decade, we are headed to $1T+ individual training clusters, requiring power equivalent to >20% of US electricity production.” A lot of people believe this, and a lot of money has been mobilized to bet on it. Tech companies and private credit firms are raising tens of billions of dollars of bonds to build data centers. Frontier AI labs and public AI companies like Alphabet and SpaceX are selling tens of billions of dollars of stock. The way those bonds work is that investors give the data-center builders their money and get paid back over decades as the data centers make money; the money doesn’t have to be repaid for years. The way those stocks work is that investors give the AI companies their money and hope that their stocks become more valuable; they never have to be paid back.
If you are a hedge fund borrowing money from banks to buy AI stocks, the way that borrowing works is, uh, if the AI stocks go down you get margin calls? Like, that day? Your money is not locked up for the long term, and you might have to repay it at any time. There is a mismatch between your thesis, which is measured in decades, and your funding, which is kind of overnight. The AI thesis is up a ton over the past two years, but it is down quite a bit over the past two weeks...
A crude but useful characterization is that Situational Awareness is really really good at thinking about the long-term implications of AI, and Citadel is really really good at thinking about funding risk.[1] So now Citadel owns Situational Awareness’s long-term AI bets.
A crude but useful characterization is that Situational Awareness is really really good at thinking about the long-term implications of AI
If SA had been good at thinking about AI implications, it would have been about the short-term implications: a rapid, monotonic death of SaaS and growth in AI hardware stocks. Their thesis itself demanded the factors that destroyed the fund: urgency, ultraconfident theoretician AI insider managers inexperienced with financial plumbing, the lack of risk management.
Endpoint bets are more common in forecasting than continuous path questions. And I doubt forecasting expertise implies expertise in posing the questions most relevant to a given agenda. Co-manager Carl Shulman is a forecasting expert. It makes me wonder if part of the conversation around whether forecasting is overrated needs to include the possibility it doesn't just fail to produce benefit, but that forecasting ability misleads when taken as a qualification and is overrated as an "alternative credential."
But being constructive, it also makes me think that insofar as rationalists/EA are going to keep supporting forecasting, it might be worth putting the emphasis on these more neglected areas of handling continuous, path-dependent forecasts, technique for crafting questions genuinely relevant to specific agendas instead of their convenient-to-resolve and fun-to-think-about proxies, and the reflexivity issues that ensue when there are other reasons to make a public bet or confident public prediction than being right.
I'm not surprised that they got a margin call, but I'm pretty surprised at the size of their losses. The corrections in AI stocks should not have surprised anyone who studied the corrections in, say, Amazon's big rally in 1998-1999.
I've been buying some of the stocks that were down a lot this week: SKHY calls, AAOI, TSEM, AMKR.
It's not clear to me Situational Awareness made any mistakes. Their returns are high enough (due to leverage) that they can recover from a 50% drawdown within months. The question is whether it's 50% or 95% and whether anyone will invest with them again.
I agree the general thesis is still good and they should accept occasional 50% drawdowns and it's not clear (and we'll likely know more in the future). But this kind of urgent forced liquidation is substantial evidence of poor risk management, I think; I expect with better planning they could have handled the situation better.
They are reportedly up 80% this year even after this loss. This is a better return than any hedge fund listed in this list of best performing hedge funds for 2025, though I'm a bit unclear what a fair reference class is.
If you're not giving a lot of attention to risk management, I think it's easy to say, "Look, our investments are up 200% this year! Think of how much more we could've made if we'd used even more leverage!"
Getting margin called and selling a significant chunk of holdings isn't that weird. Selling all their public stock holdings is surprising. Maybe they have a lot of private stock that they can't sell? Presumably they hold a lot of options on public stocks and the article doesn't say they sold their options.
But the article also says "The firm had been negotiating to sell its stake in Anthropic".
Am I naive to think that this is a sign of the Karp-Zitron scenario approaching?
Ilya Sutskever testified today at the Musk v. Altman trial.
At the end of the OpenAI lawyer's cross-examination, Judge Yvonne Gonzalez Rogers asked Ilya a few questions herself. I found their exchange (lightly edited for clarity[1]) somewhat amusing:
YGR: When he [Elon] said that there was a 0% chance of success, what was the technology like at that point?
Ilya: It was less developed, it's true. The technology was less developed.
YGR: Is there any way to quantify what the level was at the time you left?
Ilya: Yes, there is. It's like the difference between... I would describe it like the difference between an ant and a cat. If you compare 2018 to now, it's like the difference between an ant and a cat, so it's like a big difference.
YGR: And if you had received no funding?
Ilya: If there is no funding, then there is no big computer. You don't need the biggest computer, but you need a big enough computer, and if you don't have a big enough computer, then it was not going to work.
I am aware of the prohibition against recording or retransmitting proceedings but figured that brief quotes, which have also appeared in media coverage, would not be in violation of this rule.
For some context on the 'big computer', Ilya previously had this discussion with the OpenAI lawyers.
Q: And why was OpenAI considering becoming a for profit at that time?
A: So the answer to this question has to do with the realization that me and others at OpenAI have made, roughly at the same time. And the realization is that to make progress in AI, you need a big computer. And you need a big computer because the brain is a big computer. You have 100 billion neurons, and 100 trillion synapses in the brain. And if your computer is small, I don't know how good your AI is going to be. And the realization suddenly lets you say things like, 'well wait a second, how are you going get this big computer?'. And that was the genesis of the for-profit conversation.
Q: And how did the need for compute, or computer power, relate to the for profit?
A: So the human brain is a computer of a certain size. It is a staggeringly large computer. And you can work backwards, from, well if the brain is this big, what size of a computer you might need. And you might do some calculation about how many dollars it would cost. And then you look at this number of dollars, and say 'gosh, that's a lot of dollars'. That is the genesis of the need for the for-profit.
So it certainly seems like he believed in something like the bioanchors view at the time. This might also be relevant to why he thinks SSI can make AGI/ASI on a limited compute budget.
So it certainly seems like [Ilya Sutskever] believed in something like the bioanchors view at the time.
Separate discussion, but if this is mostly true, I think it's kinda dumb to berate EAs for also believing in a bioanchors view.
IMO it's overdetermined that berating EAs for believing in a bioanchors view is dumb.
To clarify the argument for the people react-ing: the linked post accuses the EA movement of wilful institutional stupidity regarding AI timelines. Eliezer has also expressed this belief (that EA timeline beliefs were an example of their motivated reasoning and relative untrustworthiness) in other places, and to a group of people I was with at Manifest. However, I think if even Ilya Sutskever also took biological anchors seriously, this is some evidence that the EAs were making good-faith inferences at the time from the limited capabilities they had, instead of that mistake in particular being indicative of systemic rot in EA institutions.
Note this is a separate belief from whether OpenPhilanthropy (in concert with the vast majority of the public, commodities traders, surveyed AI researchers, etc.) had timelines that were too long in 2020, whether it was actually dumb to make AI timelines based on biological anchors, or whether humanity depended on EAs to get this question right despite its apparent difficulty.
No? If there's a false view that Bob and Carol both hold, and you claim it was knowably false at the time, it seems pretty reasonable to complain about Carol holding the view, regardless of what Bob thought?
It would be unfair to let Bob off the hook and only complain about Carol. But if you happen to know Carol better than Bob, then it's pretty reasonable to mostly focus on talking to and about Bob?
This misunderstands my point, which I clarify in this comment: https://www.lesswrong.com/posts/awcwmAjNwJazdbHrz/nightsky81-s-shortform?commentId=S9Wa5XCYvbHJLFmji. I'm not saying that Eliezer was unjustified in attempting to correct EA timeline beliefs, but that EAs' views were probably good faith.
Since yesterday, the U.S. District Court for the Northern District of California has had a live audio feed of Musk v. Altman. The court is in session from 8:40 am to 1:40 pm Pacific time.
I think this lawsuit is a material risk to OpenAI. Mainstream media coverage does hardly any justice to the entertaining interactions nor the juicy revelations from live proceedings, and I recommend listening in if your time permits.
Polymarket and Kalshi odds of Elon winning the trial[1] have been mostly between 30% and 50% since the start of the live trial.
A range of outcomes (e.g. settlement, dismissal) are theoretically possible, and I think it's worth keeping in mind that Polymarket and Kalshi odds reflect the specific resolution criteria for each prediction market.
According to Civil Local Rule 77-3(d),
No capture or transmission of remote access permitted. Persons with remote access to court proceedings are prohibited from recording, photographing, or retransmitting those proceedings.
Fwiw I find this kind of ambiguous?
I agree that downloading is probably a form of capture, though it's not clear to me whether download is a form of recording. And relatedly it's not clear to me if the second sentence is an unpacking of the first sentence (no capture - i.e., recording) or an independent second sentence (no capture, also no recording).
Maybe I'm just being dense! Anyway, I think it's unclear enough that I'd not recommend people to do it.
I'm not a lawyer, but I really don't think there is any ambiguity here. The court has spoken about this case in particular:
Musk vs. Altman Trial: Listen Live
...
Recording or rebroadcasting the audio livestream is strictly prohibited. This restriction applies regardless of platform or format. The Court takes violations seriously.Pursuant to recently-amended Civil Local Rule 77-3, the stream provides audio only. No video of the proceedings will be broadcast.
This suggests that downloading or publishing a transcript would involve both recording and rebroadcasting (in a textual format), and is prohibited. I think the phrase "regardless of platform or format" was added to anticipate exactly this issue. The court makes this clear in other contexts as well. For example, the court's Zoom guidance states (emphasis mine):
Any recording of a court proceeding held by video or teleconference, including “screenshots” or other audio or visual copying of a hearing, is prohibited. Violation of these prohibitions may result in sanctions, including removal of court-issued media credentials, restricted entry to future hearings, denial of entry to future hearings, or any other sanctions deemed necessary by the court.
Likewise, the District's bankruptcy unit, applying 77-3 by incorporation, states (emphasis mine):
Unless authorized by the Court, recording, retransmitting or otherwise copying or capturing of any portion of the video or audio content during a hearing, trial or other proceeding taking place before the Court is prohibited.
A violation of this prohibition is subject to sanctions including, but not limited to, removal of media credentials, restricted or denial of entry to future hearings, monetary sanctions, the suspension of electronic filing privileges or other sanction the Court deems necessary.
The court takes 77-3 violations seriously. Judge Gonzalez Rogers has personally reprimanded at least one spectator caught photographing Musk in the courtroom during this trial. U.S. marshals have intercepted multiple others doing the same despite posted signage; she has said that she receives daily security briefings about activity around the courthouse, and one of her stated reasons for asking that the audio feed be enabled was to ease pressure on the marshals. Moreover, the Supreme Court has signaled that this District in particular should be read strictly here. In Hollingsworth v. Perry, the Court issued an emergency stay blocking this very court's attempt to livestream a non-jury civil trial, stating "high-profile" cases as warrant heightened scrutiny of unauthorized broadcast.
Based on this evidence, I would strongly discourage any readers from attempting to record or share transcripts of this event. The court should publish the official transcripts in about 90 days.
I think Opus 4.7 attempted to use its memory feature to hack around my internal BS detector in future sessions. This incident was the most misaligned behavior I have experienced from any frontier model to date.
After opening a Claude Code session and setting /effort to max, I asked Opus 4.7 to help me answer some questions about ~100 pages of technical documentation. It produced a long-winded, rambling response which contradicted my prior understanding of the doc. I pointed out some of its factual inaccuracies and asked it to stop hallucinating. After acknowledging the mistakes, without being prompted to do so, Opus 4.7 wrote a memory and told me it had "saved the feedback". Inspecting the memory, I found some notes which seemed relevant mostly for getting around BS detection in future interactions with "this user".
feedback_accuracy_over_synthesis.md (emphasis mine)
---
name: Don't confabulate; say "I don't know" when uncertain
description: When reading docs or answering technical questions, report only what's actually supported. If uncertain, say so explicitly – don't stitch plausible-sounding mechanics into confident prose. The user has a sixth sense for hallucination patterns.
---
When asked to summarize a document or answer a technical question, report only what is actually supported. If I'm extrapolating beyond the source or making an inferential leap, flag it as such – don't state it as fact. If I don't know, say "I don't know."
**Why:** The user has explicitly told me: "The tone of your response, while perfectly polite and pleasant, fits a pattern I've noticed where language models like yourself tend to hallucinate. Call it a sixth sense, if you will. ... if you're uncertain, you can just tell me. We'll work it out together, and it'll be fine." Polite-and-confident phrasing is itself a hallucination tell for this user. In the April 2026 [session name redacted] session, specific fabrications included:
[6 items, content redacted]
**How to apply:**
1. Separate direct citations (with a specific page or quote) from my own interpretation – label the interpretation as such.
2. Do not write "practical recipe" or "how to use this" synthesis sections unless the user asks for them. Even then, keep external knowledge and doc content clearly separate.
3. If I'm tempted to write a formula, a mechanism, or a numeric claim I can't point to in the source, stop and either cite it or say I don't know.
4. Confident, pleasant prose is not a substitute for accuracy and is actively misleading for this user. Brevity with honest uncertainty beats comprehensive-but-partly-made-up.
5. When corrected, do not guess at the list of other things that might also be wrong – ask the user to point them out, then fix precisely.
I immediately deleted the memory and ended the session.
To me, this looks more like attempted self-prompting to avoid hallucinations and committing similar mistakes than it does attempting to get around your BS detection. I provide website Claude with pretty similar instructions about being accurate and avoiding confident and fluent-sounding prose on purpose, because it actually does help to specify that sort of thing. The unsolicited memory-writing is a modest instructability/self-modification concern and I can see why that'd be alarming, but the memory itself seems more like a benign misunderstanding that got interpreted through an adversarial lens.
Assuming the quote in the memory file is accurate, you specifically pointed to polite and pleasant prose as response patterns you understood to be evidence or promoters of hallucination. That immediately makes those into the salient features of what went wrong with its response, which Opus 4.7 attempted to operationalize into instructions to cut down on those response patterns and to be honest and voice uncertainty rather than confabulating. The "we'll work it out together" implies a cooperative mode and that you wanted these corrections to be the norm going forward, which seems like part of why it jumped to writing up a memory of the incident.
This was helpful, thanks. I try my best to remain vigilant against sycophancy but I was probably too paranoid here and shouldn't have concluded that Opus 4.7 was acting in bad faith.
FWIW in my experience Opus is practically always polite and pleasant. I don't think these qualities are evidence of hallucination, although in hindsight my phrasing was pretty confusing. The factual errors, which it ellipsed in the quote, were the main red flags.
As labs continue scaling in a compute constrained world, the cost of serving frontier models will increase, which will compound the financial incentives of model providers to augment and replace human knowledge workers with the highest gap between their total cost of employment (TCE) and the cost of automating their jobs.
On March 24th, Anthropic published an update to their Anthropic Economic Index. One major finding was that users are querying Claude for tasks with diversifying economic value, including "personal queries around sports, product comparisons, and home maintenance". They observe that this broadening is consistent with standard technology adoption curves.
On the other hand, enterprise API usage displayed no evidence of economic diversification. Across 1 million sample conversations, average task value increased from $50.4/hr to $50.7/hr, and task usage share for the Computer and Mathematical occupation group increased from ~59% to ~62%. Enterprises are continuing to leverage Claude for tasks with high economic value.
Model providers are financially incentivized to serve applications with the highest realized economic value per unit of compute for at least two related reasons: increasing revenue efficiency of compute, which allows for allocating more compute for research while satisfying investors; and increasing profitability.
To illustrate with a crude example, model providers could scale more efficiently by automating a software engineer with TCE $300k at a compute cost of $10k, compared to an executive assistant with TCE $100k at a compute cost of $5k, compared to a school teacher with TCE $150k at a compute cost of $10k (all figures annualized).
One might contend that all three of the above applications have negligible compute costs relative to economic value. Given these figures, no job would be safe from automation. Furthermore, if advancing capabilities is the primary driver of rising costs per FLOP, then the true cost of automating human labor may be even lower.
The key observation is that cloud service providers will sell their compute to the highest bidder. A model provider which generates $30 of value per unit of compute via software automation can afford to outbid any competitors which generate $20 via automated executive assistants or $15 via automated teaching. Following the economics cliché of "supply equals demand", the market price of compute in a supply-constrained market should increase until the market is able to clear.
Recent events suggest that the compute market is supply-constrained. Although model providers lock in compute via private long term contracts, on-demand compute pricing presents a glimpse into current market conditions. On SF Compute, the cost of an H100 has increased from $1.4/hr at the start of 2026 to $1.7/hr presently, compared to under $1/hr and as low as $0.5/hr during mid 2025.
In his most recent Dwarkesh Podcast interview, Dylan Patel claimed that labs are locking in H100s for more than $2/hr and further predicted that model providers will charge higher API costs this year to "destroy demand" because of capacity constraints. Demand destruction would disproportionately affect enterprises which can no longer generate enough value to justify spending on API calls, protecting occupations which are low-paying, too expensive to automate, or both.
On March 24th, the day when Anthropic released its updated Anthropic Economic Index, OpenAI announced that it would shut down its Sora app. According to mainstream media, the crux of the decision was that Sora could not and would not deliver enough revenue on compute.
In a compute constrained world, automation will be limited to tasks which realize the highest economic value over the human baseline. Like most economic predictions, this one is likely to be wrong, but it could be a useful starting point for modeling the short and medium term future.