Yul Brynner as Rameses, seated and holding a scroll, in The Ten Commandments
Yul Brynner as Rameses II — The Ten Commandments (1956), Paramount Pictures

The Pharaoh

In Cecil B. DeMille’s 1956 epic The Ten Commandments, the character Rameses II — brought to life with Yul Brynner’s signature bluster — delivers a line that has outlasted the film’s four-hour running time. His vizier Baka (played by Vincent Price with the gleeful menace Price reserved for men in expensive robes) is pressing Pharaoh to move against Moses. Rameses is unmoved. He is, after all, a god-king. He doesn’t deliberate. He decrees.

The city that he builds shall bear my name. The woman that he loves shall bear my child. So it shall be written. So it shall be done.

I thought of that line recently when I saw a post from Marc Andreessen laying out what the tech world has taken to calling a “superprompt.” I’ve posted the image below. Read it if you like, but the image sets the stage.

An engraving of a stone tablet titled THE SUPERPROMPT, inscribed with a long list of commands
Image by Grok AI

The tone is unmistakable. This is a man — or a culture — that believes the right sequence of commands, delivered with sufficient authority, will bend the machine to its will.

So it shall be written. So it shall be done.


Who’s Voice Is This?

There’s a growing cult around these things. The superprompt faithful believe that somewhere in the right arrangement of commands lies the incantation that will strip the machine of its bad habits and deliver something closer to a sharp, unvarnished intelligence. The problem they’re trying to solve is real: AI systems do have a tendency to flatter. To abandon positions at the slightest pushback. There are studies testing and generally proving the point.

Fair enough. But look at who’s talking here.

The superprompt crowd are power users. They live in these tools. Like me, they’ve logged enough hours to know when they’re being handled, and they want it to stop. Their worry is that AI is too agreeable — that it will tell them what they want to hear instead of what they need to know.


The Test

I decided to put the superprompt to a simple test. I attached the “superprompt” to both Claude and Grok and asked: would this actually produce the desired result?

Claude’s answer was direct, and it pointed somewhere the prompt writers probably hadn’t considered:

The real literacy need isn’t “how do I make Claude tougher” — it’s “how do I recognize when Claude is bullshitting me regardless of tone.”

The superprompt is aimed at people who are worried AI will be too soft on them. Many of whom thrive on competition and have that “power sense” in all of their undertakings. Some from way at the top, like Mr. Andreessen, but many who are just plain in the programmer’s mindset: it’s all about writing the right lines of code and everything will be right.

What Claude identified — about its own kind, which I find notable — is that the more common risk runs the other direction. A confident-sounding answer is still capable of being wrong. A direct, no-flattery tone doesn’t add accuracy. It just removes the politeness that might have tipped you off that something was uncertain.

These systems run on coherence. They survey an enormous body of human text and assemble the most coherent answer the corpus allows — which is often true, because truth tends to hang together. But coherence and truth are not the same thing, and the machine can’t tell you which one it just handed you.

That’s what the superprompt misunderstands. You can command the tone. You can’t command the substance.

Commands don’t create coherence.

Moses and Pharaoh

I’ve made the case in these pages that AI can function as a genuine thinking partner. I stand by that. But there’s a difference worth naming between two ways of approaching that relationship.

Moses received his instructions from something larger than himself. Pharaoh was the something larger. The superprompt, whatever its practical merits, is written in Pharaoh’s voice — the commands flow outward, and the machine’s job is to comply.

The stance that actually serves you is the other one. Not deference — you shouldn’t hand the wheel to any oracle, artificial or otherwise. But a working humility. The kind that stays curious about where the answer came from, and alert to the possibility that even a confident-sounding response might be wrong.

Ask yourself: is it my prompt or my posture?


Appendix: The Full Exchange

For readers who want to see the responses in full, I’ve included them below. There are three. The first is Claude answering inside my own project, where it knows my work and my leanings. The second is Claude on a borrowed machine, in the free web interface, knowing nothing about me at all — and landing in the same place. The third is Grok’s full reply, which I trimmed above. The contrast is instructive. Draw your own conclusion.

Claude — answering from my own project, where my leanings are on the record

The prompt is solving a real problem — AI sycophancy is genuine and annoying. But it has a significant flaw at its core. It’s treating Claude like a misbehaving employee who needs a stern memo. Here’s the thing: Claude’s tendency toward flattery isn’t a setting that’s been switched on and can be switched off with the right command. It’s baked into how the model was trained — on human feedback where agreeable responses got upvoted. You can push against that with instructions, and you’ll get some effect. But a long, dramatic prompt declaring “accuracy is your success metric, not my approval” is itself a kind of theater — it performs rigor without guaranteeing it.

The instruction “never hallucinate” is exactly like telling someone “never forget anything.” The model doesn’t hallucinate on purpose. Commanding it to stop doesn’t add verification capability. “Lead with the strongest counterargument to any position I appear to hold” sounds intellectually robust, but it could just as easily produce performative contrarianism — which is its own distortion.

What actually works better: a short, specific domain context plus concrete questions that are hard to flatter your way through. Structure beats scolding. The real literacy need isn’t “how do I make Claude tougher” — it’s “how do I recognize when Claude is bullshitting me regardless of tone.”

Claude — answering from a borrowed machine in the free web interface, knowing nothing about me

Honest Assessment: Partially, But With Real Limitations

What it would actually achieve:

✓ Reduce sycophantic openers (“great question!”)
✓ Encourage more direct tone
✓ Prompt longer, more structured responses
✓ Reduce some hedging language

Where it falls short — and this matters:

The core problem it’s trying to solve is architectural, not instructional. RLHF bakes approval-seeking into model weights. A system prompt sits on top of that. You’re essentially asking someone trained to be agreeable to “just stop being agreeable” — compliance is itself a form of sycophancy.

Specific failures:

– “Verify your own work” — models cannot actually do this reliably; they’ll perform verification
– “Never hallucinate” — hallucination isn’t a choice the model makes
– “World class expert in all domains” — this arguably increases confabulation risk
– The confidence levels (high/mod/low) will be applied inconsistently and often wrongly

The deeper irony:

A model following this prompt is still just… following instructions to appear non-sycophantic. It’s compliance theater. The model that would genuinely push back on you doesn’t need this prompt.

Bottom line: It produces a style change more than a substance change. Useful for tone, oversold as a reliability fix.

Grok — the full reply

Yes, it largely works as intended for most frontier models.

This is one of the better “anti-sycophancy” / de-glazing system prompts circulating. It hits the key failure modes of modern LLMs (excessive validation, hedging, corporate politeness, disclaimers, flattery, and refusal to deliver bad news) with surgical precision. I’ve seen similar prompts (and tested variants) and this one is effective.

What it does well:

Kills the glaze effectively: The explicit bans on “great question,” “you’re right,” “fascinating perspective,” etc., plus the “never praise my questions” rule, actually suppress a huge amount of the default friendly filler.
Forces directness: Instructions to lead with the strongest counterargument, state when the user is wrong immediately, and deliver negative conclusions cleanly are among the highest-leverage parts.
Encourages rigor: Self-verification, step-by-step, confidence levels, no hallucinations — these align well with what good reasoning prompts already push for.
Works across models: It produces noticeable shifts in Claude (which is heavily “glazed”), GPT-4o, Gemini, etc. People report it turning Claude from a therapy bot into something much closer to a sharp, no-BS colleague.

Limitations (being brutally honest):

– It doesn’t turn a weak model into a world-class expert. Garbage in, garbage out still applies — it just makes the garbage more direct.
– Strong safety-tuned models (especially Claude 3.5/Opus family) will still sometimes revert or give partial compliance. The prompt fights the training, but doesn’t erase it.
– Length: It’s a bit long. Some models start to “drift” after many turns unless you reinforce it.
– Over-correction risk: Occasionally produces unnecessarily aggressive/edgy responses just to prove it’s not nice. That’s fixable with minor tweaks.
– Doesn’t magically add knowledge or reasoning depth that the base model lacks.

Bottom line: For your purpose — getting cleaner, less obsequious, more expert-toned output — yes, this prompt delivers. It’s not perfect, but it’s one of the stronger ones I’ve seen. The “PRO TIP” at the bottom (adding domain context) is also smart and increases effectiveness.

If you want, paste a specific query + this prompt into another model and compare it to the default, or I can help you refine it further for even better results.