Researchers surprised that with AI, toxicity is harder to fake than intelligence

Thank you for reading this post, don't forget to subscribe!

The next time you encounter an unusually polite reply on social media, you might want to check twice. It could be an AI model trying (and failing) to blend in with the crowd.

On Wednesday, researchers from the University of Zurich, University of Amsterdam, Duke University, and New York University released a study revealing that AI models remain easily distinguishable from humans in social media conversations, with overly friendly emotional tone serving as the most persistent giveaway. The research, which tested nine open-weight models across Twitter/X, Bluesky, and Reddit, found that classifiers developed by the researchers detected AI-generated replies with 70 to 80 percent accuracy.

The study introduces what the authors call a “computational Turing test” to assess how closely AI models approximate human language. Instead of relying on subjective human judgment about whether text sounds authentic, the framework uses automated classifiers and linguistic analysis to identify specific features that distinguish machine-generated from human-authored content.

“Even after calibration, LLM outputs remain clearly distinguishable from human text, particularly in affective tone and emotional expression,” the researchers wrote. The team, led by Nicolò Pagan at the University of Zurich, tested various optimization strategies, from simple prompting to fine-tuning, but found that deeper emotional cues persist as reliable tells that a particular text interaction online was authored by an AI chatbot rather than a human.

The toxicity tell

In the study, researchers tested nine large language models: Llama 3.1 8B, Llama 3.1 8B Instruct, Llama 3.1 70B, Mistral 7B v0.1, Mistral 7B Instruct v0.2, Qwen 2.5 7B Instruct, Gemma 3 4B Instruct, DeepSeek-R1-Distill-Llama-8B, and Apertus-8B-2509.

When prompted to generate replies to real social media posts from actual users, the AI models struggled to match the level of casual negativity and spontaneous emotional expression common in human social media posts, with toxicity scores consistently lower than authentic human replies across all three platforms.

To counter this deficiency, the researchers attempted optimization strategies (including providing writing examples and context retrieval) that reduced structural differences like sentence length or word count, but variations in emotional tone persisted. “Our comprehensive calibration tests challenge the assumption that more sophisticated optimization necessarily yields more human-like output,” the researchers concluded.

What's Hot

Who Is Trump’s New Fed Chair Pick?

X deactivates European Commission’s ad account after the company was fined €120M

Is it true that… you should take vitamin C when you’ve got a cold? | Alternative medicine

Researchers surprised that with AI, toxicity is harder to fake than intelligence

X deactivates European Commission’s ad account after the company was fined €120M

‘Kids can’t buy them anywhere’: how Pokémon cards became a stock market for millennials | Pokémon

Still Holiday Gift Shopping? Let Me Text You the Best Holiday Deals Straight to Your Phone, No Subscription Needed

Palestinians flee Gaza City districts as Israel says first stages of assault have begun

Client Challenge

The Best Weekend Resorts Near NYC

Who Is Trump’s New Fed Chair Pick?

X deactivates European Commission’s ad account after the company was fined €120M

Is it true that… you should take vitamin C when you’ve got a cold? | Alternative medicine

News

Categories

Useful links

What's Hot

Researchers surprised that with AI, toxicity is harder to fake than intelligence

The toxicity tell

Related Posts

News

Categories

Useful links

Subscribe to Updates