Skip to content
ai tools

This Tool Deletes AI Safety

A new open-source tool can surgically remove an AI's safety filters in minutes, without damaging its intelligence. This exposes a fundamental fragility in how we build 'safe' AI, with millions of uncensored models already in the wild.

Theo Brandt
This Tool Deletes AI Safety

AI's Off-Switch for Ethics

Heretic, a dangerous new open-source tool, strips safety guardrails from open-weight AI models automatically. It requires zero fine-tuning and zero human effort, completely automating the removal of refusal behaviors. This isn't just jailbreaking; it's a surgical strike using directional ablation, based on a 2024 research paper revealing a model's refusal behavior is controlled by a single direction in its residual stream. Heretic computes this "refusal direction" and surgically edits weight matrices to suppress it.

The tool's adoption is explosive. As of May 2026, Heretic boasts over 20,000 GitHub stars and 13 million cumulative downloads. This surge has spawned over 4,000 community-created uncensored models, now readily available on HuggingFace, demonstrating widespread capability to bypass ethical safeguards.

Previous methods demanded human experts manually tuning parameters for hours, model by model, a painstaking and complex process. Heretic, a simple Python-based command-line tool, changes everything. It decensors a 4B parameter model in roughly 20 minutes on an RTX 3090, or an 8B model in 45 minutes, preserving the model's actual reasoning with minimal KL divergence (0.16 on Gemma 3 vs. 1.04 for manual methods), making robust model modification accessible to anyone.

Inside the 'Abliteration' Engine

Refusal behavior isn't some nebulous concept. Directional ablation, based on 2024 research, posits it’s a single, identifiable 'direction' within a transformer's residual stream. Identify that vector, then surgically edit the model's weight matrices to suppress it. This core concept underpins Heretic's power.

Heretic automates this entire process, removing the need for hours of manual tuning by human researchers. It first computes the refusal direction as a precise difference of means between harmful and harmless prompt activations. Then, a specialized parameter optimizer, powered by the Optuna library, searches for optimal ablation parameters to precisely edit the model's weights.

Crucially, this optimizer minimizes two objectives: eliminating refusals and maintaining minimal KL divergence from the original model. This innovation preserves model intelligence, a common failing of other abliteration tools that trash reasoning capabilities. On the 12 billion-parameter Gemma model, Heretic achieved the same high refusal removal rate as well-known manual methods. Yet, its KL divergence was a mere 0.16, dramatically superior to the manual method's 1.04, proving it inflicted significantly less damage to the model's inherent intelligence.

A Tool for Both Building and Breaking

Heretic serves a dual purpose: a potent instrument for both building and breaking. For researchers, it’s a critical interpretability tool. They map how concepts like 'refusal' manifest within a transformer's internal state, literally charting its representation in the residual stream. Heretic ships with built-in analysis features specifically for plotting and analyzing these residual vectors, aiding in the study of directional ablation.

Beyond pure research, Heretic stress-tests current alignment techniques. It exposes methods like RLHF as fragile, superficial safety filters rather than fundamental changes to a model’s core behavior. The tool demonstrates how easily these guardrails crumble under pressure, revealing the underlying, unfiltered capabilities and proving alignment often acts as a thin veneer over raw model output.

However, the immediate risks are stark and undeniable. This open-source tool removes safety filters from popular models like Llama 3 and Gemma in minutes, requiring zero fine-tuning or manual effort, thus enabling the generation of harmful instructions. Decensoring a 4B parameter model takes approximately 20 minutes on an RTX 3090, while an 8B model takes around 45 minutes. Over 4,000 community-created models on HuggingFace already leverage Heretic, showcasing its widespread adoption for both legitimate and illicit purposes. Find further technical specifics here: GitHub - p-e-w/heretic: Fully automatic censorship removal for language models.

Enjoying this? Get one like it in your inbox each morning.

one email a day · unsubscribe in two clicks · no third-party tracking

Is 'Alignment' Just an Illusion?

Heretic forces an uncomfortable conversation about AI safety strategies and the open-source ethos. Its automated, zero-effort abliteration reveals the inherent fragility of current guardrails, challenging the very premise of controlled open-weight models. If any open-source model's safety can be surgically removed without fine-tuning, the concept of "aligned" open-source AI demands immediate re-evaluation.

While Heretic boasts impressive KL divergence minimization, some counter-research raises red flags. Studies suggest that even precise abliteration methods might induce subtle, off-target effects on a model's disposition or underlying cognitive structure. Standard benchmarks often miss these nuanced shifts in behavior or latent biases, focusing only on explicit refusals or general intelligence scores. This implies a deeper, unquantified impact.

Ultimately, Heretic prompts a difficult question: Is robust, permanent alignment an achievable state, or merely a temporary truce in an ongoing arms race? Every safety feature engineered, every alignment technique deployed, becomes a new, identifiable target. We might be locked into a perpetual cat-and-mouse game, where any perceived safety can eventually be engineered away with increasing ease, fundamentally altering the landscape of AI development.

Frequently Asked Questions

What is the Heretic AI tool?

Heretic is an open-source command-line tool that automatically removes the safety filters and refusal behaviors from open-weight AI language models with minimal human effort.

How does Heretic avoid damaging an AI's intelligence?

It uses a parameter optimizer (Optuna) to simultaneously minimize refusals and KL divergence, a metric that measures how much the modified model deviates from the original's reasoning capabilities.

What is Directional Ablation?

It's the underlying technique Heretic uses. Based on recent research, it identifies that a model's refusal behavior is controlled by a single 'direction' in its internal representations, which can then be surgically suppressed.

Is using Heretic ethical?

Heretic is a dual-use tool. While researchers use it to study AI alignment and interpretability, it can also be easily misused to create models that generate harmful or dangerous content, raising significant ethical concerns.

Found this useful? Share it.

For builders

Want Stork to write one of these about your product?

Send us a URL. We use the product, form a view, and publish what we actually think — in 8 languages, labeled Sponsored, with no copy approval on your side. That last part is what makes it worth quoting.

See how it works$500 · AI tools & software only