Large language models can convey empathy without appearing human and display humanness without empathy, according to research published in Nature Communications.
Bennett Kleinberg of Tilburg University led a five-study investigation comparing human writers against OpenAI systems. In the first two trials, people and GPT-4 composed relationship advice and descriptions, working either with or without instructions to sound human. External evaluators then guessed the author. Prompts to sound human raised perceived humanness ratings only for GPT-4, narrowing the gap between human and automated text.
Prompts telling models to avoid sounding like artificial intelligence produced identical patterns across GPT-4 and GPT-4o in the third study. A fifth experiment revealed that GPT-4 boosted humanness scores by adopting informal, conversational language and surface-level style cues. Human participants made few adjustments under the same instructions.
Evaluators in the fourth experiment recorded a distinct split between empathy and humanness across generated advice. The authors determined that whether readers judge text as human depends on the writing context and the presence of prompts encouraging mimicry.
Co-authors from Tilburg University, University College London, the University of Amsterdam, and the IMT School of Advanced Studies in Lucca conducted the project without competing financial interests. Nature Communications received the paper on November 5, 2024, and accepted it on August 24, 2026, publishing the work on September 3, 2026.