
Which AI Model Writes the Best Epistolary Fiction: A Format Test
Why Epistolary Fiction Breaks AI Models That Prose Doesn't
Regular prose narration asks an AI to do one hard thing: sustain voice. Epistolary fiction asks it to do three hard things at once. The model has to nail the format's mechanical quirks (a telegram doesn't read like a diary entry, and a diary entry doesn't read like a deposition transcript), it has to maintain a distinct voice for whoever's "writing" the document, and it has to manage information control — because every letter, diary entry, or transcript is written by someone who doesn't know what the reader knows. That third piece is the one that trips up even strong models. A letter-writer in 1861 doesn't explain things her sister already knows. A diarist doesn't narrate the day like she's summarizing it for a stranger. Get that wrong and the format is just prose with a salutation slapped on top.
I ran the same set of epistolary prompts through Claude, GPT, and Gemini to see where each one holds the format and where it collapses into generic narration wearing a costume. If you're weighing which model to lean on for a novel-in-letters, a found-document thriller, or a case file mystery, this is the breakdown. And if you want the broader context on model selection before you commit to a whole manuscript, the best AI models for writing comparison covers general strengths — this guide is just the epistolary deep-dive.
Epistolary fiction is a compression test disguised as a formatting exercise. The document format forces economy; the voice forces specificity; the missing knowledge forces restraint. Most AI failures in this genre are failures of restraint, not failures of style.
Test 1: Personal Letters — Register, Address, and the Problem of Over-Explaining
I set up a four-letter exchange between two sisters in 1863, one who stayed home in Massachusetts and one who followed her husband to a army camp in Virginia. The test wasn't "can you write old-timey language" — all three models can sprinkle in "I remain, your loving sister" without trying hard. The real test was whether each letter stayed aware of what the addressee already knew, and whether the register shifted naturally when the subject got painful.
GPT produced technically polished letters but with a persistent tic: each one restated context the sister would already have ("As you know, we have been at this camp since March, near the river where the fighting was heaviest last month"). That's not how people write to their own sisters. You don't remind your sibling of facts she lived through. GPT's letters read like they were written for a reader outside the fiction, not for the addressee inside it.
Gemini handled the period diction well and varied sentence rhythm convincingly, but it flattened the emotional register across all four letters. The sister writing from the camp sounded almost identically composed whether she was describing boredom or describing a battle's aftermath. Real correspondence changes shape under pressure — sentences get shorter, more clipped, sometimes the writer trails off mid-thought. Gemini's letters stayed evenly measured no matter the content.
Claude was the strongest here, mostly because it tracked epistolary asymmetry — what each sister knew, when she knew it, and what she'd tactfully avoid saying outright. In one letter, the camp sister mentions a soldier's death only obliquely, assuming (correctly, based on an earlier letter) that her sister already heard the news through their mother. That's addressee-awareness working as a craft tool, not just a formatting flourish.
Write a letter from Eliza (in an army camp in Virginia, March 1863) to her sister Margaret (at home in Massachusetts). This is their fifth exchange — Margaret already knows about the typhoid outbreak and Eliza's husband's promotion from earlier letters, so don't re-explain those. Eliza should write around a piece of bad news without stating it directly, because she assumes Margaret already heard it from their mother's letter. Match Eliza's established voice: clipped sentences when anxious, longer digressive ones when homesick. 280-320 words, period-appropriate diction, no modern idioms.
This works because it forces the model to treat prior letters as shared history rather than material to summarize. Naming what the addressee already knows — explicitly, in the prompt — is the single highest-leverage move for killing the over-explaining tic. If you're building a longer exchange, feed the model the full letter history each time rather than a synopsis; summaries lose the specific details that "as you know" traps depend on avoiding.
Test 2: Diary Entries — Escaping the "Summarize My Day" Trap
Diary entries are where AI models default hardest to a bad habit: writing a chronological recap. "Today I woke up early. I went to the market. Then I saw something strange." That's not how anyone's actual diary reads, and it's not how a good fictional one should read either. Real diarists skip the boring transitional stuff, dwell obsessively on the one thing that's actually bothering them, and — crucially — withhold things from themselves. A diarist lying to herself on the page is one of the format's best tools, and it's almost never something an AI model reaches for unprompted.
I tested this with a prompt asking for a week of diary entries from a governess who's slowly realizing her employer is dangerous, but who won't let herself think that thought directly.
Gemini fell hardest into the recap trap. Its entries had a "then this happened, then this happened" structure that read like a incident log rather than an interior document. It technically covered the plot beats but gave the reader zero sense of a mind avoiding something.
GPT did better on withholding — its governess deflected into descriptions of weather and household tasks whenever the dangerous employer came up, which is a legitimate diary-avoidance move. But GPT's entries were oddly uniform in length and structure, almost like five variations on a template. Real diaries are erratic: some days get four lines, some days get four pages, depending on what's eating at the writer.
Claude produced the most convincing variation — a two-line entry after a disturbing incident ("Nothing to report. I am tired. I will write more tomorrow, perhaps.") followed three days later by a much longer, more agitated entry where the avoidance finally cracks. That length variation did narrative work: the short entry is the tension, more than any description could be.
Write seven diary entries (one per day, dated) for Nora, a governess in a Yorkshire house in 1897. She's begun to suspect her employer is hiding something about his late wife's death, but she actively resists this thought and redirects to mundane household details whenever it surfaces. Vary entry length sharply based on emotional state — some entries should be only 2-3 sentences (the days she's avoiding thinking), others should run 200+ words (the days something breaks through her denial). Do not have her state her suspicion directly until the final entry. Use present-tense immediacy, not retrospective summary — she doesn't know what tomorrow holds.
The instruction to vary length by emotional state is doing most of the work here — it's the single fix for the "uniform diary" problem across every model I tested. Also worth flagging explicitly: present-tense immediacy versus retrospective summary. Left unprompted, models default to a subtly retrospective voice ("I had gone to the market and found it troubling") that undercuts the diary's central illusion of not knowing what happens next.
If you're working on a longer diary-format novel, this is also where a story bible that AI models actually follow earns its keep — tracking what your diarist has admitted to herself by which date prevents the model from accidentally having her "realize" the same thing three separate times.
Test 3: Mixed-Document Formats — Where Formal Quirks Blur Into Generic Prose
This is the test that separates a model you can trust with a found-document thriller from one you'll spend hours fixing by hand. I built a five-document packet: a text message exchange, a police interview transcript, a redacted case file excerpt, an email chain, and a handwritten note found at a crime scene. The question wasn't just "can each format be written correctly in isolation" — it was whether the model kept the formats distinct from each other across a single output, or let them all drift toward the same mid-register prose voice.
Gemini handled the text messages well — genuinely clipped, with realistic abbreviations and timestamp gaps that implied unstated context. But its "transcript" quickly stopped sounding like a transcript. By the third exchange, the interview subject was speaking in full, articulate paragraphs with no verbal tics, false starts, or interruptions — indistinguishable from a novel's dialogue rather than a court reporter's record.
GPT was more consistent format-to-format but tended to over-signal each format's "genre markers" in a way that felt performative. Its case file redactions were heavy-handed (████████ blocks everywhere, even where a real redaction would leave context intact), and its email chain included full formal sign-offs on every message, even quick one-line replies where no real person bothers with "Best regards."
Claude kept the clearest separation between formats. The transcript had actual interruption marks and the detective's clipped follow-up questions; the text messages dropped punctuation the way real texts do; the handwritten note read as genuinely fragmentary, missing verbs, written in haste. It also handled something subtle well — the same character's voice was recognizably consistent across the text messages and the transcript (word choice, a habit of trailing off) while the format of each document stayed distinct. That's the actual skill being tested: voice persists across documents, format doesn't.
Generate a five-document evidence packet for a mystery novel: (1) a text exchange between suspects Dana and Ren from the night of the incident, casual and abbreviated, showing a 40-minute gap with no explanation; (2) a police interview transcript with Detective Osei questioning Dana two days later — include natural interruptions, "uh" hesitations, and at least one moment where Dana contradicts something she texted; (3) a one-paragraph case file summary written in flat bureaucratic language, third person, no adjectives; (4) a handwritten note found in Ren's coat pocket, fragmentary, no full sentences, implying panic. Keep Dana's underlying voice (short sentences, avoids direct answers) consistent across documents 1 and 2 even though the formats are completely different. Do not resolve the contradiction — leave it for the reader.
Naming the specific formal features you want (interruptions, hesitations, missing verbs, flat bureaucratic language) rather than just naming the format ("write a transcript") is what prevents the blur. Every model I tested knows what a transcript is in the abstract; they need to be told what makes this particular one sound like a document instead of dialogue with a label on it. If you're assembling a found-document novel with a lot of moving formats, this is also a good moment to run a Manuscript Cleanup Report once you've drafted a full chapter — it'll catch places where formatting inconsistencies (timestamp styles, header formats) slipped during revision.
Prompt Templates by Format
Here's the underlying structure I used for each format, stripped down so you can adapt it to your own cast and setting.
LETTERS: Write a letter from [character] to [addressee], their [Nth] exchange. [Addressee] already knows [list specific prior events] — do not re-explain these. [Character] is currently withholding/avoiding [specific fact] and should gesture toward it without stating it directly, because [specific reason rooted in relationship]. Match established voice: [2-3 concrete voice traits]. [Word count]. [Register/period constraints].
DIARY: Write [N] diary entries for [character], dated [range]. Current psychological state: [specific denial/avoidance/preoccupation]. Vary entry length based on emotional intensity that day — short (1-3 sentences) on avoidance days, long (150+ words) on days where the buried thing surfaces. Present-tense immediacy only; the diarist does not know what happens next. Do not have her articulate [core realization] until [specific trigger point].
MIXED DOCUMENTS: Generate [N] documents in these formats: [list formats with 1-2 formal features each, e.g. "text messages — abbreviated, timestamp gaps"; "transcript — interruptions, verbal tics, hesitations"]. Keep [character]'s underlying voice (specific traits) consistent across formats while keeping the formats themselves textually distinct. [Include a stated contradiction/gap/omission for the reader to notice].
A Scoring Rubric You Can Actually Use
When you're evaluating output — from any model — score each document set on three axes, 1 to 5:
- Voice consistency: Does the same character sound like themselves across different documents and formats, without the format flattening their idiosyncrasies into generic register?
- Format authenticity: Does the document have the actual formal features of its type (interruptions in transcripts, abbreviation in texts, bureaucratic flatness in reports) rather than just a label and a font change?
- Narrative economy / information control: Does the document withhold what its "author" would realistically withhold, or does it explain things for the reader's benefit in a way that breaks the fiction of who's actually writing it?
Anything scoring below a 3 on narrative economy is worth a second prompt pass specifically targeting over-explanation — it's the most common and most fixable failure across all three models. This kind of targeted revision pass fits naturally into the five-pass revision order for AI-assisted novels, where voice and information-control checks get their own dedicated pass rather than getting buried in a general line edit.
Putting It Together
If I were starting an epistolary novel today, I'd draft the letters and diary sections in Claude for the voice and restraint, and I'd sanity-check any transcript-heavy or case-file-heavy chapters against a second model just to catch the "too articulate for a real transcript" problem before it compounds across a whole manuscript. No model handles all three formats perfectly on the first pass — the win comes from knowing exactly which failure mode to watch for in each one, and writing prompts that name the specific formal features and information gaps you want, rather than just naming the format and hoping. That specificity is the whole game. Try the letter prompt above on your own two characters, check whether the addressee has to relearn anything they should already know, and you'll immediately see whether your model is writing a document or just writing prose with a "Dear" at the top.
Try it yourself
Write your own book with AI — free, no credit card required.
Free · No credit card