
How to Use Constraint Stacking to Force AI Prose Compression
AI prose has a weight problem. Feed a model a bloated scene and ask it to "tighten this up," and you'll get back something roughly the same size, maybe with a few adverbs trimmed. Ask it to rewrite from scratch and you'll often get something longer — more scene-setting, more interiority, more of everything you didn't need. This isn't a bug exactly. It's how these models are trained: to be helpful, to be thorough, to fill the space.
The fix isn't a better one-line instruction. It's layered, sequential pressure — what I call constraint stacking. You build a prompt that attacks the prose from three different angles at once: how sentences are built, what information is actually necessary, and what emotional register the scene needs to inhabit. When all three layers are active simultaneously, the model has almost no room to bloat. It has to compress. And compression, done right, doesn't strip meaning — it concentrates it.
This technique is especially useful if you're trying to write a book with AI at a commercial pace without sacrificing prose quality. You can generate fast and then compress hard, rather than babysitting every sentence in real time.
Why AI Defaults to Expansion (And What Constraint Stacking Fixes)
Watch what happens when you paste a 400-word scene into most AI tools and write "make this tighter." You get 380 words back. The model interprets "tighter" through its training — which associates quality writing with completeness, with showing rather than telling, with fully rendered detail. It's not wrong, exactly. But it's optimizing for a different goal than yours.
The deeper issue is that vague instructions leave the model's defaults intact. "Tighter" doesn't override the tendency to add a beat of reflection after each action. "More concise" doesn't stop it from expanding a single sensory detail into a full sentence when it thinks that detail matters. You're giving it a direction without giving it a wall to hit.
Constraints aren't restrictions on creativity. They're the walls that make the ball bounce in the right direction.
This is why single-instruction prompts plateau. The model hears "cut 20%" and trims some adjectives. It doesn't restructure. It doesn't question whether a thought belongs in this scene at all. Constraint stacking forces that deeper interrogation by making the model satisfy multiple competing demands at once — and when those demands conflict, compression is usually the only solution that satisfies all of them.
If you're already thinking about revision at scale, it's worth reading The Five-Pass Revision Order for AI-Assisted Novels, which places compression work inside a larger editorial sequence. Constraint stacking fits cleanly into that framework, usually around passes two and three.
The Three Constraint Layers: Syntactic, Semantic, and Tonal
Layer One: Syntactic Constraints
Syntactic constraints operate at the sentence level. They control structure, length, and rhythm. Examples: maximum sentence length, prohibition on subordinate clauses, no passive constructions, no sentences that begin with "It was" or "There were." These are mechanical rules the model can actually track.
Useful syntactic constraints for compression include:
- Hard sentence cap: "No sentence longer than 14 words."
- Clause prohibition: "No subordinate clauses. Main clause only."
- Verb primacy: "Every sentence must have a strong active verb as its main predicate."
- No gerund openers: "Do not begin any sentence with a gerund phrase."
These feel almost arbitrary, but that's the point. When the model can't write "She stood there, thinking about what he'd said and wondering whether she'd made a mistake," it has to choose: the action or the thought. That forced choice is compression.
Layer Two: Semantic Constraints
Semantic constraints control what information is allowed in the prose at all. This is the harder layer because it requires judgment, not just rule-following. But you can approximate it with specific prohibitions.
Good semantic constraints:
- "Remove all interiority that restates information the reader already has."
- "No character can explain their emotion — show one physical action instead."
- "Cut any sentence that describes what a character is about to do rather than doing it."
- "Each paragraph must advance either the plot or the character's understanding of their situation. Not both. Not neither."
That last one is brutal in the best way. A lot of AI-generated prose advances nothing — it decorates. Semantic constraints starve the decoration.
Layer Three: Tonal Constraints
Tonal constraints set the emotional register and prohibit tonal drift. AI prose loves to soften edges, add warmth, explain feelings. Tonal constraints push back against that tendency by naming the specific emotional texture you want and banning its opposite.
Examples:
- "Tone: cold and clinical. No warmth, no hope, no comfort."
- "The prose must feel like controlled rage — not grief, not nostalgia."
- "Do not use any word or phrase that softens the scene's violence."
Tonal constraints interact with syntactic ones in interesting ways. "Cold and clinical" plus "no sentences longer than 12 words" produces something very different from "cold and clinical" with no syntactic constraint at all. The stack creates a specific texture you couldn't get from either layer alone.
How to Stack Constraints Without Losing Narrative Substance
Here's where writers worry: won't stacking this many rules strip the life out of the prose? It can, if you stack blindly. The key is knowing what each constraint is targeting and making sure you're not deleting load-bearing structure by accident.
Before you stack, identify what the scene actually needs to do. Not what it currently does — what it must do. Does the reader need to feel the protagonist's hesitation? Does the scene need to convey that this location is hostile without saying so? Write that down. Then apply constraints that compress everything else while leaving those essential functions intact.
A practical approach: run constraints in sequence rather than all at once if the scene is complex. Apply syntactic constraints first and read the output. Then add semantic constraints and run again. Then layer in tonal constraints. This gives you control over what each layer is doing, and you can catch it when a constraint accidentally kills something that mattered.
You can use the Manuscript Diff tool to compare versions side by side — it's particularly useful at this stage for spotting when compression went too far and stripped out a detail that was actually earning its place.
One rule I've found reliable: always preserve one specific sensory anchor per scene. One concrete, physical detail the reader can grab. Constraints can eat everything else, but that anchor keeps the scene from feeling like a synopsis. Specify it explicitly in your prompt: "Preserve the detail about the glass on the windowsill. Everything else is negotiable."
If you're working with complex genre worlds — fantasy settings, magic systems, future tech — compression gets trickier because you're carrying worldbuilding weight. The guide on Fantasy Magic Systems: Constraints AI Will Respect has a useful section on how to compress exposition without losing reader orientation. The same principles apply to any spec-fic scene with load-bearing detail.
Prompt Examples: From Bloated Draft to Compressed Scene
These are real-pattern prompts you can adapt directly. Each one stacks all three constraint layers — adjust the specifics to match your scene.
Rewrite the scene below using all of the following constraints simultaneously:
SYNTACTIC: No sentence may exceed 12 words. No subordinate clauses. Active voice only.
SEMANTIC: Cut all interiority that restates what the reader already knows. Remove any sentence that describes what a character is about to do instead of doing it. Each paragraph must move the scene forward in time or change a character's situation.
TONAL: Flat, exhausted, post-crisis. No hope, no resolution, no warmth. The character is past feeling — she's just moving.
Preserve: the image of the broken lock on the back door. Everything else is subject to cuts.
[PASTE SCENE HERE]
This works because the 12-word cap makes it nearly impossible to nest emotions into action beats. The semantic rule about "about to do" eliminates a huge category of AI padding. The tonal flatness prevents the model from adding a hopeful beat at the end, which it will almost always try to do otherwise. Tweak the word cap up to 16 if you want slightly more room for rhythm.
Take the following dialogue scene and compress it by at least 40% using these layered rules:
SYNTACTIC: Dialogue lines may not exceed 8 words each. No dialogue tag more complex than "said" or "asked." No action beats that occur mid-dialogue line — action beats appear only between lines.
SEMANTIC: Remove any line of dialogue where a character explains what they're feeling. Remove any line that is answered by a line saying essentially the same thing in different words. Keep only lines that change the power dynamic between the two characters.
TONAL: Dangerous and careful. Both characters are hiding something. The prose should feel like two people walking on ice.
[PASTE DIALOGUE SCENE HERE]
The 8-word dialogue cap is aggressive — it forces the model to find the sharpest version of every line. The semantic rule about "power dynamic" gives the model a structural test to apply: if a line doesn't shift who has leverage, cut it. The "walking on ice" tonal image is the kind of concrete metaphor models respond to better than abstract descriptors like "tense." Try swapping it for your scene's specific dynamic — "two people pretending to be friends," "someone saying goodbye without saying it."
Compress this interior monologue passage. Constraints:
SYNTACTIC: Alternate between sentences of 5 words or fewer and sentences of 10–15 words. No more than two consecutive sentences at the same length. No questions — the character is not wondering, she's knowing.
SEMANTIC: Every sentence must either reveal new information about the character's situation OR advance her decision. Cut anything that circles back to something already established. No metaphors for her emotional state — describe physical sensation only.
TONAL: Controlled fury. Precise, not scattered. She is certain, not spiraling.
[PASTE INTERIOR MONOLOGUE HERE]
The alternating sentence rhythm is unusual but powerful — it produces a kind of pulse that reads as controlled intensity without you having to name it. The "no questions" constraint is a small thing that has an outsized effect: question sentences in interiority slow pace and suggest uncertainty. If your character is certain, ban the question mark entirely. Adapt the decision anchor to whatever your character is actually deciding.
When to Release Constraints and Let the Prose Breathe
Constraint stacking is a scalpel, not a default setting. There are scenes that need compression and scenes that need space, and knowing the difference is as important as knowing how to stack.
The obvious moments to release constraints: grief, revelation, the slow unfolding of something the reader has been waiting for. These scenes earn their length. Compressing a death scene into 12-word sentences might work for one genre (crime fiction where everything is clipped and hard) and destroy another (literary fiction where the prose itself is mourning). Genre matters here. If you're writing a romance novel with AI, the compression technique that works for your thriller-paced chapter breaks will be completely wrong for your emotional climax.
The goal of constraint stacking isn't short prose. It's prose where every word is load-bearing. Sometimes load-bearing prose is long.
A useful test: after you've compressed a scene, read it aloud and notice where you instinctively want to slow down. If the prose won't let you — if the short sentences keep pushing you forward past a moment that needed weight — you've over-compressed. Add back one syntactic constraint release ("one sentence in this paragraph may exceed 20 words") and run again.
You should also think about what compression does to your reader's pacing across a full manuscript. A chapter of relentlessly compressed prose followed by another compressed chapter creates a flatline effect. If you're deep in revision and thinking about the Beta Reader Workflow for AI-Assisted Manuscripts, consider asking beta readers specifically about pacing rhythm — they'll often identify over-compression before you do, because they're reading fresh.
The technique also has natural limits for scenes that need to carry a clue network across your novel. When a scene is doing structural work — planting information the reader needs to notice but not quite register — aggressive compression can make that detail disappear. The clue that worked when it was embedded in three sentences of surrounding prose might vanish entirely when you cut to the bone. In those cases, compress everything except the carrier sentence, and tell the model exactly which sentence to protect.
One last thing: track which constraint combinations produce results you actually like. When you find a syntactic + semantic + tonal stack that generates prose in your voice, save it. That's your compression template. Drop it into your Series Bible Template as a style note, and you can apply it consistently across a multi-book project without rebuilding the logic every time.
The real payoff of constraint stacking isn't any single tighter scene. It's that over time, you develop an instinct for which layer is causing the problem in a bloated passage. You start reading your AI drafts and thinking: "This needs a semantic constraint — it's repeating information." Or: "The tonal layer is wrong — it keeps softening." That diagnostic sense is worth more than any individual prompt, because it means you're editing AI prose with the same clarity you'd bring to your own.
Start with one scene you know is overwritten. Apply the three layers one at a time, reading the output after each pass. By the third pass, you'll understand — physically, in your reading body — what each layer is doing. That's the fastest way in.
Try it yourself
Write your own book with AI — free, no credit card required.
Free · No credit card