Can listeners distinguish between LLM-generated Baroque music and the real thing? I wanted to find out, so I created a blind listening quiz. It features 16 one-minute excerpts, mixing works by human composers with music generated by GPT-6 Astra and Claude Opus 5.5.
The more responses the quiz gets, the more meaningful the data will be, so once you’ve taken it, feel free to share the quiz with anyone who might be interested. AI optimists, skeptics, and haters are all welcome.
After the quiz closes on October 12, I’ll analyze the responses, highlight interesting findings, and publish the results. Until then, please don’t discuss or reveal individual answers.
About the Quiz
All excerpts were rendered from MIDI files to help level the playing field between the AI and human competitors. This removes performance and recording quality from the equation, requiring listeners to judge the music itself.
AI Track Generation & Selection
Each AI piece was generated from a single prompt, and each prompt was used only once. Every output made it into the final quiz, with one exception — a piece I posted on YouTube recently was originally intended for the quiz, but I liked it so much I decided to share it on social media instead. Some might call this cherry-picking, but if it is, it made the AI pieces worse on average.
Human Track Selection
When choosing pieces by human composers, I aimed for a wide variety of forms and national traditions (French, German, Italian, etc.). My main criterion was that a piece had to be obscure enough that the average listener wouldn’t recognize it, but notable enough to have been recorded. (Most of the pieces have been recorded extensively.)
The downside to using lesser-known works is that the music representing Team Humanity is lower quality than it would be if I had stuck to the greatest hits. But I don’t consider this much of a design flaw, as no one is seriously comparing AI to the best of Bach, only to an average Baroque piece.
Nonetheless, I’d say the competition is stacked in humanity’s favor. On one hand, you have the first results from hastily written prompts, produced by LLMs still in their digital diapers. On the other, you have the music of professional composers whose work has lasted hundreds of years. Telling them apart should be easy, right?
Baroque Focus
The quiz focuses on the Baroque style because it’s one of the few styles LLMs currently excel at. A quiz involving Classical or Romantic pieces would be too easy. But I expect LLMs to improve in those styles soon, so we’ll probably be doing this again next year with Beethoven and Brahms.
Additional Data
The quiz ends with a few questions about musical training and listening experience. This should tell us whether trained musicians and frequent listeners are better at detecting AI-generated music.
Along with the results, I’ll publish the prompts, models, and effort settings used for each AI-generated track.
I found it very limiting to have to give a verdict on each one before listening to the next. With Scott Alexander's Art Turing Test, I could study all the pictures while considering my verdicts. Would it be possible to put all the pieces and their questions on one page?
So I only listened to the first, and I found I couldn't make a judgement. Without saying anything specific about that piece, I found it unlistenable because of the hideously synthetic performance. Anything would sound like AI performed like that. I appreciate you may not be able to hire musicians to perform them, but I just can't get past it.
Can listeners distinguish between LLM-generated Baroque music and the real thing? I wanted to find out, so I created a blind listening quiz. It features 16 one-minute excerpts, mixing works by human composers with music generated by GPT-6 Astra and Claude Opus 5.5.
Link: AI Music Detection Quiz
Duration: 18–20 minutes
No sign-up or email required
The more responses the quiz gets, the more meaningful the data will be, so once you’ve taken it, feel free to share the quiz with anyone who might be interested. AI optimists, skeptics, and haters are all welcome.
After the quiz closes on October 12, I’ll analyze the responses, highlight interesting findings, and publish the results. Until then, please don’t discuss or reveal individual answers.
About the Quiz
All excerpts were rendered from MIDI files to help level the playing field between the AI and human competitors. This removes performance and recording quality from the equation, requiring listeners to judge the music itself.
AI Track Generation & Selection
Each AI piece was generated from a single prompt, and each prompt was used only once. Every output made it into the final quiz, with one exception — a piece I posted on YouTube recently was originally intended for the quiz, but I liked it so much I decided to share it on social media instead. Some might call this cherry-picking, but if it is, it made the AI pieces worse on average.
Human Track Selection
When choosing pieces by human composers, I aimed for a wide variety of forms and national traditions (French, German, Italian, etc.). My main criterion was that a piece had to be obscure enough that the average listener wouldn’t recognize it, but notable enough to have been recorded. (Most of the pieces have been recorded extensively.)
The downside to using lesser-known works is that the music representing Team Humanity is lower quality than it would be if I had stuck to the greatest hits. But I don’t consider this much of a design flaw, as no one is seriously comparing AI to the best of Bach, only to an average Baroque piece.
Nonetheless, I’d say the competition is stacked in humanity’s favor. On one hand, you have the first results from hastily written prompts, produced by LLMs still in their digital diapers. On the other, you have the music of professional composers whose work has lasted hundreds of years. Telling them apart should be easy, right?
Baroque Focus
The quiz focuses on the Baroque style because it’s one of the few styles LLMs currently excel at. A quiz involving Classical or Romantic pieces would be too easy. But I expect LLMs to improve in those styles soon, so we’ll probably be doing this again next year with Beethoven and Brahms.
Additional Data
The quiz ends with a few questions about musical training and listening experience. This should tell us whether trained musicians and frequent listeners are better at detecting AI-generated music.
Along with the results, I’ll publish the prompts, models, and effort settings used for each AI-generated track.
Take the quiz: AI Music Detection Quiz