"Genlangs" and Zipf's Law: Do languages generated by ChatGPT statistically look human?

Justin-Diamond

"Genlangs" and Zipf's Law: Do languages generated by ChatGPT statistically look human?

by Justin-Diamond

1 min read31st Jan 20242 comments

2

Data ScienceLanguage & LinguisticsMachine Learning (ML)Probability & StatisticsAI

Frontpage

This is a linkpost for https://arxiv.org/abs/2304.12191

OpenAI's GPT-4 is a Large Language Model (LLM) that can generate coherent constructed languages, or "conlangs," which we propose be called "genlangs" when generated by Artificial Intelligence (AI). The genlangs created by ChatGPT for this research (Voxphera, Vivenzia, and Lumivoxa) each have unique features, appear facially coherent, and plausibly “translate” into English.

This study investigates whether genlangs created by ChatGPT follow Zipf's law. Zipf’s law approximately holds across all natural and artificially constructed human languages. According to Zipf’s law, the word frequencies in a text corpus are inversely proportional to their rank in the frequency table. This means that the most frequent word appears about twice as often as the second most frequent word, three times as often as the third most frequent word, and so on. We hypothesize that Zipf’s law will hold for genlangs because (1) genlangs created by ChatGPT fundamentally operate in the same way as human language with respect to the semantic usefulness of certain tokens, and (2) ChatGPT has been trained on a corpora of text that includes many different languages, all of which exhibit Zipf's law to varying degrees. Through statistical linguistics, we aim to understand if LLM-based languages statistically look human. Our findings indicate that genlangs adhere closely to Zipf’s law, supporting the hypothesis that genlangs created by ChatGPT exhibit similar statistical properties to natural and artificial human languages.