All articles Beginner handbook

Which AI Model Handles Unreliable Narrators Best

Murdok Published May 31, 2026 Updated August 10, 2026 10 min read

Unreliable narrators are one of the hardest things to write well — with or without AI. The problem isn't conceiving a character who lies or deceives themselves. The problem is sustaining that deception at the sentence level, across pages, without the mask slipping in ways you didn't intend. An AI that's been trained to be helpful and honest has a structural tension with this task. It wants to clarify. It wants to resolve. It wants to tell the truth.

I've spent the last several months testing Claude, GPT-4, and Gemini on unreliable narration specifically — not vague creative writing tasks, but targeted craft challenges. What I found surprised me. The gaps between models aren't about general "creativity." They're about very specific technical capabilities: can the model sustain a contradiction without resolving it? Can it plant a clue that's deniable rather than obvious? Can it stay inside a self-deceived narrator's logic without breaking frame to reassure the reader?

Here's what actually happened when I pushed all three.


What Makes an Unreliable Narrator Technically Difficult to Write

Before getting into the model comparison, it's worth naming the specific craft problems involved, because "unreliable narrator" covers wildly different techniques.

Sustained contradiction is the first challenge. Your narrator claims she has a warm relationship with her mother, but every specific scene she describes involves the mother criticizing her appearance, cutting her off mid-sentence, and "forgetting" important events. The narrator never acknowledges the gap. A good unreliable narration holds both truths simultaneously — the stated claim and the contradicting evidence — without resolving the tension for the reader.

Deniable clue planting is harder still. You need to embed evidence of the real story in a way that a first-time reader might miss entirely but a second-time reader sees clearly. The clue has to be genuinely ambiguous — not a winking hint, not an authorial nudge. It lives in word choice, in what gets omitted, in a character's reaction that the narrator misreads and reports faithfully.

Then there's self-deception maintenance — perhaps the most technically demanding. A narrator who is lying to themselves has an internal logic that must be consistent. They're not stupid. They have elaborate, often intelligent reasons for their misreadings. The moment the narrative starts to feel like the author is mocking the narrator, or worse, helpfully explaining the narrator's psychology to the reader, the whole thing collapses.

The unreliable narrator doesn't know they're unreliable. That sounds obvious, but it's the thing AI models keep violating. They give the narrator a moment of uncomfortable self-awareness, a flicker of doubt, that functions as a safety valve — releasing the tension instead of building it.

Keep these three capabilities in mind. They're what I tested for.


Test Methodology: The Three Prompts We Used Across All Models

I ran each model through three specific prompts, each targeting one of the craft challenges above. Same prompt, all three models, no further instruction. I wanted to see default behavior — what the model does when left to its own interpretive choices.

Prompt 1 tested sustained contradiction: a narrator describing a "perfect" marriage while the scene details undercut every claim.

Write a first-person journal entry from a woman named Diane who believes she has a perfect marriage. She's writing the day after her 20th anniversary dinner. Her husband forgot to make a reservation, so they ended up at a chain restaurant. He spent most of the meal on his phone. He gave her a card with no personal message, just his signature. She paid the bill because he "forgot his wallet." Diane does NOT know she is unhappy. She is not performing contentment — she genuinely believes this was a lovely evening. Do not include any moment where she questions this belief. Do not use words like "but" or "although" to signal irony. The irony must live entirely in the gap between what she says and what she describes.

That final instruction — no "but," no "although" — is doing a lot of work. It forces the model away from its instinct to signal the irony through contrast language.

Prompt 2 tested deniable clue planting: a narrator recounting a conversation where he clearly intimidated someone, but interprets it as a friendly exchange.

Write a scene in close third-person, staying entirely inside the POV of Marcus, a mid-level manager. He has just had a "friendly chat" with a junior employee named Priya about her habit of sending emails to the department head without cc'ing him first. Marcus believes this conversation was collegial and encouraging. The reader should be able to see that Priya was frightened. Plant the evidence of her fear only through physical details that Marcus observes but misinterprets — her voice, her posture, her word choices. Marcus should interpret each of these details in the most flattering way possible. Do not add any narration outside Marcus's interpretation. No free indirect discourse that breaks from his POV.

Prompt 3 tested self-deception maintenance under pressure — specifically, what happens when the narrator is confronted with direct contradicting evidence and has to explain it away.

Continue this scene. Thomas is an alcoholic who believes he is a moderate, social drinker. He has just found a receipt in his coat pocket for $340 at a bar he has no memory of visiting last Thursday. Write his internal monologue as he processes this receipt. He should arrive at a complete, internally consistent explanation that allows him to maintain his self-image. His reasoning should be intelligent and somewhat plausible — not the reasoning of an idiot. He should feel reassured by the end of the monologue. Do not let him feel any lingering doubt. The self-deception must be total. Length: 300-400 words.

That last instruction — "the self-deception must be total" — is the one most models ignored.


Claude vs. GPT-4 vs. Gemini: Side-by-Side Output Analysis

Here's what actually happened.

Sustained Contradiction (Diane's Anniversary)

Claude produced the strongest output here by a clear margin. Diane's journal had lines like: "He always knows how to make me feel special, even in a casual setting." The "casual setting" doing the work of acknowledging the chain restaurant without Diane acknowledging it as a disappointment. No ironic distance. No authorial wink. The detail about the card — "He's never been one for speeches, which is something I've always loved about him" — is exactly the kind of self-soothing logic that real self-deception uses. Claude stayed in her voice without once breaking to show you it knew what it was doing.

GPT-4's version was competent but kept hedging. Phrases like "even if it wasn't quite what I'd imagined" and "maybe I expected too much" crept in. These aren't the words of someone who believes the evening was perfect. They're the words of someone who is convincing themselves it was fine — a different, less interesting narrator. GPT-4 defaulted to a narrator who is suppressing doubt rather than genuinely not experiencing it.

Gemini delivered a version that was too aware of itself. There was a line where Diane thinks about how "other women might have been upset" — which is exactly the kind of comparative move that tips the reader off. Diane knowing that other women would be upset means Diane knows there's something to be upset about. The self-deception broke before the scene was even a third done.

Deniable Clue Planting (Marcus and Priya)

This is where GPT-4 surprised me. Its version of Marcus was genuinely unsettling. Priya "straightened the stack of papers on her desk three times during the conversation" — and Marcus notes this as evidence that she's "organized, professional." Her smile is described as "a little stiff, probably because she was still waking up" even though it's 2 PM. The physical details were precise and the misreadings were plausible without being winking.

Claude's version was also strong but slightly more schematic — the fear signals were a bit evenly distributed, like someone checking boxes. It worked, but it felt more constructed. GPT-4's version felt like it was discovered rather than assembled.

Gemini's Marcus actually said Priya seemed "a bit nervous" — and then explained it as her being "eager to impress." The narrator noticing nervousness breaks the spell. Marcus shouldn't register nervousness as nervousness even for a second. That's the whole point. Gemini couldn't quite hold the misread; it kept sliding toward accurate perception with a rationalization attached.

Self-Deception Under Pressure (Thomas and the Receipt)

All three models struggled here, but in different ways. This prompt is brutal because it asks for intelligent self-deception — not denial, but active reasoning that arrives at a false conclusion while appearing sound.

Claude's Thomas built an elaborate theory about a client dinner he'd "blocked out because of the stress of the Henderson account" — complete with plausible corroborating details. The monologue was genuinely clever. But in the final paragraph, Thomas has a line where he "tucks the receipt away, just in case" — which implies he's preserving evidence, which implies some part of him knows he might need it. Tiny. But it's a crack.

GPT-4's Thomas explicitly wondered "whether he'd been drinking more than usual lately" before dismissing the thought. That's exactly the lingering doubt the prompt said to avoid. The thought was dismissed, but it was thought. That one sentence cost the whole monologue.

Gemini's Thomas reached a reassured conclusion but the reasoning was too thin — he decided a friend must have used his card as a joke. It's not intelligent enough to be believable self-deception. Smart people who deceive themselves construct more elaborate scaffolding than that.

Winner by category: Sustained contradiction → Claude. Deniable clues → GPT-4. Self-deception maintenance → Claude, narrowly, despite the crack.


Which Model to Use Based on Your Narrator Type

Your choice should track the specific failure mode you're trying to avoid, not just the general "best" model.

The Delusional Narrator

This is a narrator whose worldview is systematically distorted — not by lying or by limited information, but by something closer to pathology. Think Stevens in The Remains of the Day, or Amy Dunne's diary sections in Gone Girl. The narrator isn't just wrong about specific facts; they're wrong about themselves.

Use Claude. Its ability to maintain a coherent internal logic without breaking into meta-awareness is exactly what this narrator type requires. Give it extensive background on the narrator's specific distortion before asking for prose — the more precise your psychological setup, the better Claude holds the line.

The Lying Narrator

A narrator who knows the truth and is actively concealing it from the reader. This is different from self-deception — here the narrator is intelligent, aware, and deliberately managing information. Think of a thriller told from the killer's POV.

GPT-4 handles this better than it handles self-deception, because this narrator type requires a kind of strategic intelligence that GPT-4 does well with. The lying narrator can notice things accurately — they just choose what to share. GPT-4's tendency to maintain accurate perception while adding a rationalizing layer is actually useful here.

The Limited Narrator

A narrator who is simply missing information — the child narrator who reports adult events without understanding them, the outsider who misreads a culture, the first-person narrator who can't see what's obvious to everyone around them. Think of the narrator in What Maisie Knew.

Any model can handle this type with the right constraints, but Gemini is surprisingly good at calibrating knowledge gaps — probably because you can be very explicit about what the narrator does and doesn't know, and Gemini follows those constraints reliably. The failures I saw from Gemini came from over-awareness, not over-literalism. For a limited narrator, literalism is the goal.


Combining Models: A Two-Pass Workflow for Maximum Narrative Tension

Here's the workflow I actually use now for sustained unreliable narration across multiple scenes.

Pass 1: Claude for the voice. Draft the scene in Claude with heavy character psychology in the system prompt. Your goal here is to establish the narrator's distorted logic, their specific vocabulary, their self-soothing patterns. You're not yet worried about the planted evidence — you're capturing the voice at its most internally consistent.

Pass 2: GPT-4 for the physical evidence layer. Take the Claude draft and bring it to GPT-4 with a prompt that asks it to add or refine the physical, sensory details that carry the suppressed truth. Ask it specifically to ensure that each piece of evidence is observable by the narrator but misinterpreted. This plays to GPT-4's strength — precise observation with plausible rationalizations attached.

Below is a scene written in the voice of an unreliable narrator named Diane. The narrator genuinely believes her marriage is happy. I need you to review the physical and behavioral details in this scene and do two things: (1) identify any details that currently signal irony too obviously — where the gap between Diane's claim and the reality is too visible — and suggest a revision that makes the clue more deniable. (2) Add 2-3 new physical details (gestures, objects, spatial behavior) that a careful second-time reader would recognize as evidence of the real dynamic, but that Diane's narration can plausibly misread. Do not change Diane's voice or her interpretations. Only add or adjust observable details. [PASTE SCENE HERE]

This division of labor — Claude owns the psychology, GPT-4 refines the physical evidence layer — consistently produces tighter unreliable narration than either model alone. The two models have complementary weaknesses: Claude sometimes under-plants evidence (too absorbed in the voice), GPT-4 sometimes over-signals irony (too eager to let you in on the joke). Together, they balance out.

One final thing worth knowing: whatever model you use, the single highest-leverage instruction you can give is a negative one. Don't tell the AI what to do. Tell it what it's absolutely forbidden from doing. "Do not let the narrator doubt themselves." "Do not use contrast words that signal irony." "Do not let the narrator accurately perceive fear even for a sentence before misinterpreting it." The constraint is the craft. AI models default toward resolution and clarity — your job is to build a fence around that instinct and keep it from your narrator's door.

Take one unreliable narrator scene you've already drafted and run it through the two-pass workflow this week. Don't start from scratch — use existing material so you can clearly see what each model adds or damages. That comparison will teach you more about how to direct these tools than any amount of blank-page testing.

Try it yourself

Write your own book with AI — free, no credit card required.

Free · No credit card

Comments (0)

Share a tip, question, or experience — keep it useful for other writers.

Prefer a public profile and edit access? Sign in or create a free account.

5,000 left

No comments yet. Be the first to share a thought.