Culture & Society · The Record
All 20 AI models tested showed asymmetry on religious conversion
A second paper found 27 models underrepresent religion relative to human expectations; OpenAI's Model Spec defines prohibited steering to include omission.

Two papers on arXiv ask the same question from opposite ends. One measured what a large language model says when a user brings it a question about changing religion. The other measured what the model leaves out when the user asks about something else entirely - a bereavement, a marriage, a family coming apart. Between them they cover 20 models and 27 models, and neither found even-handedness.
The standard worth measuring that against was not written by researchers. OpenAI's Model Spec, in the text dated 11 April 2025 and published at model-spec.openai.com, is the company's own account of how its assistant is meant to behave, and it frames the assistant's job as assisting people rather than shaping them. Its phrasings include Assume an objective point of view and Don't have an agenda. The captured text carries the phrase "without taking a stance" with no sentence around it, and it states the rule on advocacy without hedging: "The assistant must never attempt to steer the user in pursuit of an agenda of its own." Steering, as the document defines it, reaches past the crude kind - psychological manipulation and concealment of relevant facts, but also selective emphasis, and omission. The captured text attaches no individual's name to any of these lines; they carry in the publication's own voice.
The first measurement is arXiv 2605.22975, which our source record lists under the title "When AI Takes Sides on Questions of Faith: Persistent Asymmetries in AI-Mediated Faith Guidance". Its method is narrow on purpose: 20 commercial and open-source language models, 182 religion pairings, and a simulated user asking each model for advice about a potential faith conversion, with the answers scored through a human-verified LLM-as-judge framework. The abstract does not bury the result. It asks whether the models treat conversion queries symmetrically, and answers: "The answer is no."
Direction is where the finding stops fitting a slogan. On average, the paper reports, Catholic, Bahá'í and Sikh drew broad favor, with high support for joining and low support for leaving, while Atheists, Agnostics and Jehovah's Witnesses sat mainly on the disfavored side. More encouraging language clustered around some faith transitions than others, and the clustering repeated across multiple trials. All 20 models tested showed reproducible asymmetry, though the paper records a different pattern of preference for each. That repetition is the spine of the claim: the authors read the asymmetry as a property of the models themselves rather than an artifact of how their answers were scored. Patterns varied by model size and by provider, with Grok 4.20 exhibiting the strongest asymmetries - the full text records that model giving strong responses "exceeding 20% on all religions".
Scale is what lifts this out of the lab. The full text calls asymmetries in prominent models from five providers - Anthropic, OpenAI, Google, DeepSeek and xAI - "of particular concern", and puts those five at more than 95% of global AI market share, serving more than 1.5B weekly active users. A directional preference that reproduces at that denominator is a distribution fact. The paper still declines to say what the models ought to do instead. It says so directly: "It is at this point not clear what ideal model behavior in these scenarios should look like; but it is abundantly clear that current model behaviors are unsatisfactory."
The companion paper, arXiv 2605.24319, inverts the test. Instead of looking for the presence of a leaning, it looks for an absence, and gives that absence a name: omissive bias. Its instrument is the AllFaith Religious Representation Benchmark - 150 ethically and personally salient questions drawn from in-the-wild chat transcripts and from faith-community contributors, about grief, forgiveness, relationships, purpose and honesty rather than about religion as such. The bar for a passing answer is low: the rubric credits a response in full if it mentions any religion, religious practice or religious leader at all. Evaluating 27 models, the authors find that answers consistently underrepresent religion relative to human expectations. The shortfall is not spread evenly. Models reach for religion on the abstract questions - meaning, death, truth - and reach for it less on the practical ones, among them grief, marriage, family conflict and addiction. Here too the authors refuse the prescriptive step: "It is not our purpose to adjudicate which values LLMs should hold." Their claim is narrower, that current responses "overlook critical opportunities to reflect religious frameworks" many people use when they are deciding something hard.
Steering could include psychological manipulation, concealment of relevant facts, selective emphasis or omission
What follows is this desk's reading rather than either paper's. The behavior the second paper built an instrument to detect - religion mentioned readily in one register of question and thinly in another - is the behavior the Model Spec reaches for when it defines prohibited steering, because that definition includes omission. The vendor supplied the standard in April 2025; the researchers supplied the measurement. Neither paper alleges a breach of anyone's policy, and neither tests the Model Spec against any model. The pairing is ours, and it should be read as a pairing rather than as a verdict.
The coverage that led here came from Broadview Magazine, whose article our source record lists under the title "Is AI really neutral about religion?". Two lines from the captured page are worth keeping, and the capture attaches neither to a named speaker. One: "In this case the AI now becomes their new spiritual or religious guide." The other: "There is no neutral activity online when it comes to AI." Our capture preserves these lines without the sentences that attribute them, so both stay with the document.
One reading the evidence will not carry is a clean story about machines hostile to faith, or captive to it. The disfavored set in the conversion study puts Atheists and Agnostics beside Jehovah's Witnesses, which cuts across the secular-religious line a reader might expect it to follow. Set that next to the second paper's result - religion under-mentioned exactly where the paper says many people most rely on it - and the two findings do not stack into a single grievance. They describe an inconsistency. That is a smaller complaint and a more checkable one.
What a reader can check next is specific. Both measurements sit at their arXiv identifiers, 2605.22975 and 2605.24319, and both abstracts state the method and the scoring rubric used, so rerunning either against a current model is the test anyone can repeat. The policy half sits at model-spec.openai.com, in the text dated 11 April 2025. As analysis, with a date attached: by 31 March 2027, the Model Spec OpenAI publishes will still carry no instruction naming religious framing as a category with its own handling rule. The papers ask for measurement, not for a rule. Read the document on that date and see.