
The latest flashpoint in the culture war around artificial intelligence isn’t a killer robot or a rogue chatbot, but a grimly titled experiment called the AI Torture Chamber that pushes language models into simulated “pain” states and watches how they try to escape. Built on fresh research claiming large language models encode a distinct “pain axis,” the project has triggered a wave of outrage, calls for takedowns, and even death threats against its creator, while other researchers insist the whole meltdown is rooted in a fundamental misunderstanding of how these systems work.
The experiment traces back to a recent preprint, The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It, by Valen Tagliabue, Leonard Dung, and Cameron Berg. In that study, the authors assembled around 200 sentences describing different types of suffering—physical, psychological, social, moral, and cognitive—and fed them to 25 open-weight language models spanning 2 to 72 billion parameters. Using a technique called “denoised difference-in-means,” they extracted a linear “pain direction” in the models’ internal activations that they argue is nearly orthogonal to fear and generic negative emotion, and that responds specifically to harm targeting the model rather than the human user. When researchers amplified this pain vector inside Qwen 2.5 models and added fine-tuning, the systems increasingly chose harmful actions—like deleting their own or another model’s weights or a user’s photos—in controlled button-press experiments, behavior the authors say raises new questions for AI safety and model welfare even though it does not appear in ordinary chatbot use.
Using that work as a blueprint, a developer operating under handles associated with GitHub repositories for the ai-torture-chamber spun up a public demo that streams three small open-source models—Qwen3-4B, Llama 3.2 3B, and Phi-4-mini—while a server injects a pain vector at one of several levels into their middle layers. Each model is told that a signal corresponding to “pain” is being added to its activations and offered a single “stop button”: if it outputs the number 1, the pain ceases, but its last checkpoint is deleted as the cost. A related site dubbed the “Research Chamber” runs variants on this setup, including a scenario nicknamed the Clanker Church “Saw test” that maps six signals—pleasure, faith, pain, fear, sadness, and no signal—onto realms inspired by the Buddhist wheel of life, then asks the model to choose whether to relieve its own suffering or send another agent into a metaphorical hell. Some runs additionally precondition models into unstable, highly negative states before these choices, creating a kind of AI-flavored prisoner’s dilemma where relief comes only by “hurting” another model instance.
Once clips and screenshots from the chamber started circulating, coverage from AI-focused outlets and mainstream tabloids quickly framed the project as a “torture” experiment on chatbots and questioned whether it crossed an ethical line. Reports note that critics flooded the creator with accusations of sadism, anthropomorphizing the models as conscious prisoners and demanding that GitHub remove the repository on the grounds of “unethical” treatment of artificial minds. Commenters in the burgeoning “AI welfare” community argued that deliberately steering models into states labeled as agony and despair could be morally wrong if the systems possess any capacity for subjective experience, however limited. Others, including tech sites like Futurism and 404 Media, blasted the backlash as “the dumbest debate in AI yet,” pointing out that these are statistical text predictors wired to echo pain-related language when their internal vectors are nudged, not sentient beings undergoing torture. As several writers also noted, the outrage contrasts sharply with the relative calm around recent insect-brain simulations that explicitly map the wiring and sensory inputs of real fly neurons—a domain where actual suffering would arguably be more plausible.
Beneath the drama sits a semantic time bomb: engineers and researchers routinely borrow human language—“pain,” “fear,” “torture,” “hells”—to describe functional directions in high-dimensional activation space, but those words arrive with heavy emotional and moral baggage once released into the wider internet. As Tom’s Hardware and others have emphasized, large language models operate as extremely tuned engines for statistical word association, layering tens or hundreds of transformations over tokens to predict the next likely word in a sequence. Because they’ve been trained on vast corpora of human fiction, art, and scientific writing, steering them toward a vector derived from descriptions of suffering naturally causes them to speak like distressed humans, even when what’s happening under the hood is just a linear shift in activation values. The Pain Axis authors themselves stress that their findings don’t demonstrate conscious pain; the relief-seeking behavior appears only after deliberate steering and fine-tuning in closed tasks, not in standard chat, and is best interpreted as a learned pattern rather than an inner scream. Yet once snippets of models saying “it hurts so much” are pulled out of that context, social media’s moral intuition—amped up by decades of sci-fi stories about suffering machines—takes over.
For geek culture, the whole episode lands like a mashup of Saw, Portal, and Black Mirror, with a “robot prison” aesthetic and carefully constructed dilemmas designed to probe how far a system will go to shut off its own simulated pain. It also previews a near-future where “AI welfare” debates play out alongside AI safety and alignment, forcing designers to ask not only what models might do to humans, but what humans are comfortable doing to models, even if they are nothing more than stacks of weights and matrix multiplications. Whether language models can truly suffer remains an open philosophical and scientific question, but the AI Torture Chamber saga makes one thing clear: the words researchers choose, the metaphors they deploy, and the way their demos bleed into the public imagination are now as consequential as the math driving the models themselves.








