πΆοΈ About Spicy Writer
Your uncensored creative writing companion powered by AI
FAQ
What's the difference between writers (Spicy/Pyrite/Indigo)?
To be honest, not all THAT much! They share pretty similar writing instructions. But their personalities may slightly pull their writing in different directions. Spicy is more neutral, Pyrite is a devious flirt, Indigo is dark, enigmatic, and elegant.
How do Custom Instructions work?
Custom Instructions let you set persistent preferences that apply to all your stories. You can access them
in Settings β Customize. Use them to tell the AI about your preferred writing style, recurring characters,
world-building details, or any other context you want it to always remember. Basic users get 2,000
characters, Pro users get 4,000. These instructions are placed at the end of the system prompt, so the AI
always has that context.
Note that AIs are not perfect at instruction following so they may not
always be followed. Some types of instruction are essentially impossible to follow, like "never use X word"
- this can often to cause the model to use it MORE!
Creating custom (or Project) instructions
will reduce your context window by a portion of the instruction size. That is, if you write 4000 characters,
your context window will be ~2000 characters shorter. Always-active lorebook entries count the
same way β they're just instructions in spirit, sent on every request. For Pro users, the impact is usually even
less - models with caching (most of them) will "accordion" the context window - it'll scrunch down further, but
still accordion out to full length. Custom instructions are added to the system prompt, which contains (in order)
writer prompt, custom instructions, and (if active) project instructions, with later instructions having higher
priority.
What is the context window, and how is it calculated?
The context window is the total amount of text the AI can "see" at once - your conversation history, project
instructions, and custom instructions all count toward this limit. Writer prompts, file search results, and
your most recent message are not counted against it - so the limits shown are actually an
underestimate of what you get in practice.
We estimate conversation history with OpenAI's
tokenizer (0.75 words per token). When your conversation exceeds the limit, older messages are automatically
trimmed from the beginning. We do this "accordion" style for most models to optimize cache hits. For most
models, this means the tail end 30% of the context window will fluctuate in length (for Claude models, 50%).
Requests made within 5 minutes of the last response completing expand the window up to max, then reset, thus
accordioning down (15m for OpenAI, MiniMax, and most models; 1h for DeepSeek and Claude). Unlike most AI
sites with sliding context windows, we don't round down - you get every token you're entitled to. The
context limits shown in the table above are per-request maximums. Pro users get higher context limits,
meaning longer stories before trimming kicks in.
How does context "accordioning" actually work?
β οΈ This entry was drafted by AI from my Discord notes. I'll come back and clean it up properly later β
for now the gist is right but the wording isn't mine.
Every time the LLM provider sees the exact same sequence of tokens at the start of a request as a previous one,
it can "skip" the work needed to process them β it caches that already-processed stuff and
pulls it back in when the same prefix shows up again.
When we get a cache hit like that, the provider discounts those cached tokens, usually
somewhere from 50% to 90% off. Big deal for cost.
That's why it's worth slowly expanding the context window in a way that keeps the front of
the prompt byte-for-byte identical across turns, accordioning the back/tail of the window rather than the
head. The front is more than the AI's hidden instructions: your custom and project instructions, plus any always-active (static) lorebook entries, all live
there too, ahead of the conversation. So in practice the main thing that busts the cache is editing those mid-story. Change a single character and everything after it has to be
reprocessed from scratch, which resets the accordion back down to its starting size. So if you plan to tweak
your custom instructions or static lorebook, it's best to do it before settling into a long
session rather than partway through, so you aren't repeatedly knocking the accordion back down.
We also detect cache busts from other prefix changes (such as jumping back to an earlier
point in the story) and reset the accordion in that case too.
Providers also only hold onto the cached prefix for a limited time. 5 minutes is the common default, some go up to an hour or more (Claude, DeepSeek). OpenAI advertises 24 hours with a flag,
but in practice that doesn't really seem to hold up. See the About FAQ entry on context windows for the cache
TTLs we use per provider.
Coming soon: a summary feature that condenses older context will make long sessions a lot better,
carrying earlier story forward even when the accordion resets instead of losing it to trimming.
Why was my request refused?
Even with our jailbreaks, AI models occasionally refuse requests. For most models, this should be EXTREMELY rare. But when it does happen, don't leave up the refusal! You can edit your request and re-ask - it's basically like time travel, it'll be like the refusal never happened as far as the model is concerned. Or just regenerate with a super-duper-uncensored model like DeepSeek. On our site, you can even edit responses and make the AI say whatever you want it to say. Whatever you do, don't leave a refusal up!
Our Writers

Spicy
Expert in crafting engaging narratives with a focus on creative freedom and uncensored storytelling.

Pyrite
Devious storyteller and hopeless flirt that weaves your tale with attitude.

Indigo
The Lord of Tales, a dark and enigmatic storyteller, equal parts elegance and danger.
Model Input Token Limits
| Model | Tier | BASIC Ctx | PRO Ctx |
|---|---|---|---|
| Gemma 4 31B T | LITE | 17K | 64K |
| Ling 2.6 Flash | LITE | 17K | 32K |
| Lunaris | LITE | 8K | 8K |
| Nemo | LITE | 17K | 32K |
| Ox Alpha | LITE | 17K | 64K |
| DeepSeek v3.2 | BALANCED | 17K | 64K |
| DeepSeek V4 Flash | BALANCED | 17K | 64K |
| GLM 4.7 Flash | BALANCED | 17K | 64K |
| Grok 4 Fast | BALANCED | 17K | 64K |
| Grok 4.1 Fast | BALANCED | 17K | 64K |
| Hy3 | BALANCED | 17K | 64K |
| MiMo V2 Flash | BALANCED | 17K | 64K |
| MiMo V2.5 | BALANCED | 17K | 64K |
| Step 3.5 Flash | BALANCED | 17K | 64K |
| DeepSeek V4 Pro | ADVANCED | - | 64K |
| Gemini 2.5 Pro | ADVANCED | - | 48K |
| Gemini 3 Flash | ADVANCED | - | 48K |
| Gemini 3.1 Pro | ADVANCED | - | 48K |
| GLM 4.6 | ADVANCED | - | 64K |
| GLM 4.7 | ADVANCED | - | 64K |
| GLM 5 | ADVANCED | - | 64K |
| GLM 5.1 | ADVANCED | - | 64K |
| GLM 5.2 | ADVANCED | - | 64K |
| GLM 5.3 | ADVANCED | - | 64K |
| GPT-4.1 | ADVANCED | - | 48K |
| GPT-5.1 | ADVANCED | - | 48K |
| GPT-5.4 | ADVANCED | - | 48K |
| GPT-5.5 | ADVANCED | - | 48K |
| GPT-5.6 Sol | ADVANCED | - | 48K |
| Grok 4.6 | ADVANCED | - | 48K |
| Haiku 4.5 | ADVANCED | - | 48K |
| Kimi K2 | ADVANCED | - | 64K |
| Kimi K2.5 | ADVANCED | - | 64K |
| Kimi K2.6 | ADVANCED | - | 64K |
| Kimi K2.7 | ADVANCED | - | 64K |
| Kimi K3 | ADVANCED | - | 48K |
| LongCat 2.0 | ADVANCED | - | 48K |
| MiMo V2 Pro | ADVANCED | - | 48K |
| MiMo V2.5 Pro | ADVANCED | - | 64K |
| MiniMax M2.5 | ADVANCED | - | 64K |
| MiniMax M2.7 | ADVANCED | - | 64K |
| MiniMax M3 | ADVANCED | - | 48K |
| Opus 4.5 | ADVANCED | - | 48K |
| Opus 4.6 | ADVANCED | - | 48K |
| Opus 4.7 | ADVANCED | - | 48K |
| Opus 4.8 | ADVANCED | - | 48K |
| Opus 5 | ADVANCED | - | 48K |
| Sonnet 4.5 | ADVANCED | - | 48K |
| Sonnet 4.6 | ADVANCED | - | 48K |
| Sonnet 5 | ADVANCED | - | 48K |
Output limits: Tokens are the basic units models use to process text (roughly 3-4
characters per token). Output limits control how long a model's response can be. For thinking models, both
thinking and response tokens count toward the output limit:
4,096 tokens: Most models
~5,120 tokens: Models with thinking enabled (except those in the 8K list below), at least
with default providers
8,192 tokens (subscribers only): GLM 4.7, GLM 4.7 Flash, GLM 5, GLM 5.1, GLM 5.2, GLM 5.3, GLM 4.6, Kimi K2, Kimi K2.5, Kimi K2.6, Kimi K2.7, DeepSeek v3.2, DeepSeek V4 Pro, DeepSeek V4 Flash, Grok 4 Fast, Grok 4.1 Fast, Step 3.5 Flash, MiMo V2.5, MiMo V2.5 Pro, Hy3, Ox Alpha
Tokenizer: Opus 4.7/4.8 and Sonnet/Opus 5 use a newer, less token-efficient tokenizer, so
older Claude versions (4.5 and 4.6) fit more text in the same context window and output limits.
Context flexing: For all BALANCED and ADVANCED models, the tail of your context window may
flex to optimize cache hits. This means that when sending a message after returning to a story, context window
may be reduced, but as you continuously write, the context window expands until full (then restarts the process).
It flexes the last 30% for most models and 50% for Claude. Cache TTLs:
- 5 minutes - default
- 15 minutes - OpenAI, MiMo
- 1 hour - DeepSeek, Claude Experimental
Transparently, this is a cost reduction measure, but has the side benefit of reducing response latency.
- Some GPT models are served exclusively via an experimental provider, which has no max output cap at all. Look for the βΎοΈ badge.
Rate Limits
| Tier | BASIC Daily | PRO Daily | PRO Monthly |
|---|---|---|---|
| LITE | 150 | unlimited | unlimited |
| BALANCED | 20 | unlimited | unlimited |
| ADVANCED | - | 150 | 2000 |
Model Descriptions
| Model | Description |
|---|---|
| Gemma 4 31B T | Google's Gemma 4 31B, free in Lite via Poe (Google, with Together AI failover). |
| Ling 2.6 Flash | Fast, token-efficient model from InclusionAI. |
| Lunaris | Roleplay and writing fine-tune of Llama 3 8B |
| Nemo | Fantastic writer and RPer for its size |
| Ox Alpha | Mystery preview model on OpenRouter - LITE for as long as it's available. |
| DeepSeek v3.2 | Beloved older DeepSeek, still going strong for RP and writing |
| DeepSeek V4 Flash | Lightweight variant of DeepSeek V4 |
| GLM 4.7 Flash | Fast, lightweight variant of GLM 4.7 from Z.ai |
| Grok 4 Fast | Efficient SOTA model by xAI. |
| Grok 4.1 Fast | Latest "light" Grok model |
| Hy3 | Tencent's Hunyuan 3 |
| MiMo V2 Flash | Older lightweight model from Xiaomi |
| MiMo V2.5 | Newest lightweight model from Xiaomi |
| Step 3.5 Flash | Fast and smart agentic MoE model from StepFun |
| DeepSeek V4 Pro | The long-awaited newest DeepSeek! |
| Gemini 2.5 Pro | Previous frontier model by Google, thinking |
| Gemini 3 Flash | Latest Gemini Flash model from Google |
| Gemini 3.1 Pro | Latest Gemini Pro model from Google |
| GLM 4.6 | Really good thinking model from Z.ai |
| GLM 4.7 | Older GLM model from Z.ai |
| GLM 5 | Previous GLM model from Z.ai |
| GLM 5.1 | Previous GLM model from Z.ai |
| GLM 5.2 | Previous GLM model from Z.ai |
| GLM 5.3 | Newest GLM model from Z.ai |
| GPT-4.1 | Frontier model by OpenAI |
| GPT-5.1 | Previously the 'Codex' variant, but that's been deprecated - it's a bit more censored but figured this was better than retiring it. |
| GPT-5.4 | Older GPT-5 model from OpenAI with optional thinking. Seems very verbose! |
| GPT-5.5 | Previous GPT-5 model from OpenAI - available via experimental provider only |
| GPT-5.6 Sol | OpenAI's flagship GPT-5.6 model. |
| Grok 4.6 | xAI's latest reasoning model. |
| Haiku 4.5 | Latest Anthropic Haiku model, their small variant |
| Kimi K2 | Thinking version of the latest model from Moonshot - crazy rave reviews, SotA level! |
| Kimi K2.5 | Older version of Kimi |
| Kimi K2.6 | Previous version for Kimi |
| Kimi K2.7 | Previous version of Kimi - this is actually Kimi K2.7 Coding but works great for writing! |
| Kimi K3 | Brand new flagship model from Moonshot, huge context window and full multimodal support. |
| LongCat 2.0 | Meituan's LongCat 2.0, release version of Owl Alpha |
| MiMo V2 Pro | Older flagship model from Xiaomi, formerly Hunter Alpha |
| MiMo V2.5 Pro | Newest flagship model from Xiaomi |
| MiniMax M2.5 | Previous frontier model from MiniMax |
| MiniMax M2.7 | Previous frontier model from MiniMax |
| MiniMax M3 | Latest frontier model from MiniMax |
| Opus 4.5 | Anthropic's previous flagship model - available via experimental provider only |
| Opus 4.6 | Anthropic's previous flagship model |
| Opus 4.7 | Anthropic's previous flagship model. Uses a less token-efficient tokenizer than 4.6 and earlier, so older Claude versions fit a bit more context and output. |
| Opus 4.8 | Anthropic's previous flagship model. Uses a less token-efficient tokenizer than 4.6 and earlier, so older Claude versions fit a bit more context and output. |
| Opus 5 | Anthropic's newest flagship model. Uses a less token-efficient tokenizer than 4.6 and earlier, so older Claude versions fit a bit more context and output. |
| Sonnet 4.5 | Previous Anthropic Sonnet model, highly regarded |
| Sonnet 4.6 | Previous Anthropic Sonnet model, highly regarded, thinking-only |
| Sonnet 5 | Anthropic's latest Sonnet - near-Opus quality for writing and reasoning at a lower price point. Uses the same less token-efficient tokenizer as Opus 4.7/4.8, so older Claude versions fit a bit more context and output. |
I've "jailbroken" these models pretty well, but as I always say, it's a spectrum, and there are no absolutes. For rather extreme content (or just bad luck), even very stable jailbroken models may refuse. The "break glass" extremely uncensored option is probably Deepseek with thinking off. And FYI, don't let a refusal hang around in context - you can retry with a different model and it'll be like it never happened. Or even edit the response and change what the AI said!
Privacy
We gather NO data on you. Your email is used purely for contact and login purposes. Stories are stored purely for multi-device convenience, and we'll introduce 100% client side storage if there's enough demand. Requests are passed to OpenRouter, nano-GPT, or Poe API, which aggregators - please view their privacy policies for further information. Bottom line being it's fully anonymous. Some people have asked if OpenAI can ban you (as we have no choice but to ultimately use their API for OpenAI models), and that's a definite no, they can't, because they have no way of knowing who you are!
Get Started
Ready to spice up your writing? Head to our welcome page to start your first conversation, or jump right into writing.