🌢️ About Spicy Writer

Your uncensored creative writing companion powered by AI

FAQ

What's the difference between writers (Spicy/Pyrite/Indigo)?

To be honest, not all THAT much! They share pretty similar writing instructions. But their personalities may slightly pull their writing in different directions. Spicy is more neutral, Pyrite is a devious flirt, Indigo is dark, enigmatic, and elegant.

How do Custom Instructions work?

Custom Instructions let you set persistent preferences that apply to all your stories. You can access them in Settings β†’ Customize. Use them to tell the AI about your preferred writing style, recurring characters, world-building details, or any other context you want it to always remember. Basic users get 2,000 characters, Pro users get 4,000. These instructions are placed at the end of the system prompt, so the AI always has that context.

Note that AIs are not perfect at instruction following so they may not always be followed. Some types of instruction are essentially impossible to follow, like "never use X word" - this can often to cause the model to use it MORE!

Creating custom (or Project) instructions will reduce your context window by a portion of the instruction size. That is, if you write 4000 characters, your context window will be ~2000 characters shorter. Always-active lorebook entries count the same way β€” they're just instructions in spirit, sent on every request. For Pro users, the impact is usually even less - models with caching (most of them) will "accordion" the context window - it'll scrunch down further, but still accordion out to full length. Custom instructions are added to the system prompt, which contains (in order) writer prompt, custom instructions, and (if active) project instructions, with later instructions having higher priority.

What is the context window, and how is it calculated?

The context window is the total amount of text the AI can "see" at once - your conversation history, project instructions, and custom instructions all count toward this limit. Writer prompts, file search results, and your most recent message are not counted against it - so the limits shown are actually an underestimate of what you get in practice.

We estimate conversation history with OpenAI's tokenizer (0.75 words per token). When your conversation exceeds the limit, older messages are automatically trimmed from the beginning. We do this "accordion" style for most models to optimize cache hits. For most models, this means the tail end 30% of the context window will fluctuate in length (for Claude models, 50%). Requests made within 5 minutes of the last response completing expand the window up to max, then reset, thus accordioning down (15m for OpenAI, MiniMax, and most models; 1h for DeepSeek and Claude). Unlike most AI sites with sliding context windows, we don't round down - you get every token you're entitled to. The context limits shown in the table above are per-request maximums. Pro users get higher context limits, meaning longer stories before trimming kicks in.

How does context "accordioning" actually work?

⚠️ This entry was drafted by AI from my Discord notes. I'll come back and clean it up properly later β€” for now the gist is right but the wording isn't mine.

Every time the LLM provider sees the exact same sequence of tokens at the start of a request as a previous one, it can "skip" the work needed to process them β€” it caches that already-processed stuff and pulls it back in when the same prefix shows up again.

When we get a cache hit like that, the provider discounts those cached tokens, usually somewhere from 50% to 90% off. Big deal for cost.

That's why it's worth slowly expanding the context window in a way that keeps the front of the prompt byte-for-byte identical across turns, accordioning the back/tail of the window rather than the head. The front is more than the AI's hidden instructions: your custom and project instructions, plus any always-active (static) lorebook entries, all live there too, ahead of the conversation. So in practice the main thing that busts the cache is editing those mid-story. Change a single character and everything after it has to be reprocessed from scratch, which resets the accordion back down to its starting size. So if you plan to tweak your custom instructions or static lorebook, it's best to do it before settling into a long session rather than partway through, so you aren't repeatedly knocking the accordion back down.

We also detect cache busts from other prefix changes (such as jumping back to an earlier point in the story) and reset the accordion in that case too.

Providers also only hold onto the cached prefix for a limited time. 5 minutes is the common default, some go up to an hour or more (Claude, DeepSeek). OpenAI advertises 24 hours with a flag, but in practice that doesn't really seem to hold up. See the About FAQ entry on context windows for the cache TTLs we use per provider.

Coming soon: a summary feature that condenses older context will make long sessions a lot better, carrying earlier story forward even when the accordion resets instead of losing it to trimming.

Why was my request refused?

Even with our jailbreaks, AI models occasionally refuse requests. For most models, this should be EXTREMELY rare. But when it does happen, don't leave up the refusal! You can edit your request and re-ask - it's basically like time travel, it'll be like the refusal never happened as far as the model is concerned. Or just regenerate with a super-duper-uncensored model like DeepSeek. On our site, you can even edit responses and make the AI say whatever you want it to say. Whatever you do, don't leave a refusal up!

Our Writers

Spicy

Spicy

Expert in crafting engaging narratives with a focus on creative freedom and uncensored storytelling.

Pyrite

Pyrite

Devious storyteller and hopeless flirt that weaves your tale with attitude.

Indigo

Indigo

The Lord of Tales, a dark and enigmatic storyteller, equal parts elegance and danger.

Model Input Token Limits

ModelTierBASIC CtxPRO Ctx
Gemma 4 31B TLITE17K64K
Ling 2.6 FlashLITE17K32K
LunarisLITE8K8K
NemoLITE17K32K
Ox AlphaLITE17K64K
DeepSeek v3.2BALANCED17K64K
DeepSeek V4 FlashBALANCED17K64K
GLM 4.7 FlashBALANCED17K64K
Grok 4 FastBALANCED17K64K
Grok 4.1 FastBALANCED17K64K
Hy3BALANCED17K64K
MiMo V2 FlashBALANCED17K64K
MiMo V2.5BALANCED17K64K
Step 3.5 FlashBALANCED17K64K
DeepSeek V4 ProADVANCED-64K
Gemini 2.5 ProADVANCED-48K
Gemini 3 FlashADVANCED-48K
Gemini 3.1 ProADVANCED-48K
GLM 4.6ADVANCED-64K
GLM 4.7ADVANCED-64K
GLM 5ADVANCED-64K
GLM 5.1ADVANCED-64K
GLM 5.2ADVANCED-64K
GLM 5.3ADVANCED-64K
GPT-4.1ADVANCED-48K
GPT-5.1ADVANCED-48K
GPT-5.4ADVANCED-48K
GPT-5.5ADVANCED-48K
GPT-5.6 SolADVANCED-48K
Grok 4.6ADVANCED-48K
Haiku 4.5ADVANCED-48K
Kimi K2ADVANCED-64K
Kimi K2.5ADVANCED-64K
Kimi K2.6ADVANCED-64K
Kimi K2.7ADVANCED-64K
Kimi K3ADVANCED-48K
LongCat 2.0ADVANCED-48K
MiMo V2 ProADVANCED-48K
MiMo V2.5 ProADVANCED-64K
MiniMax M2.5ADVANCED-64K
MiniMax M2.7ADVANCED-64K
MiniMax M3ADVANCED-48K
Opus 4.5ADVANCED-48K
Opus 4.6ADVANCED-48K
Opus 4.7ADVANCED-48K
Opus 4.8ADVANCED-48K
Opus 5ADVANCED-48K
Sonnet 4.5ADVANCED-48K
Sonnet 4.6ADVANCED-48K
Sonnet 5ADVANCED-48K

Output limits: Tokens are the basic units models use to process text (roughly 3-4 characters per token). Output limits control how long a model's response can be. For thinking models, both thinking and response tokens count toward the output limit:
4,096 tokens: Most models
~5,120 tokens: Models with thinking enabled (except those in the 8K list below), at least with default providers
8,192 tokens (subscribers only): GLM 4.7, GLM 4.7 Flash, GLM 5, GLM 5.1, GLM 5.2, GLM 5.3, GLM 4.6, Kimi K2, Kimi K2.5, Kimi K2.6, Kimi K2.7, DeepSeek v3.2, DeepSeek V4 Pro, DeepSeek V4 Flash, Grok 4 Fast, Grok 4.1 Fast, Step 3.5 Flash, MiMo V2.5, MiMo V2.5 Pro, Hy3, Ox Alpha

Tokenizer: Opus 4.7/4.8 and Sonnet/Opus 5 use a newer, less token-efficient tokenizer, so older Claude versions (4.5 and 4.6) fit more text in the same context window and output limits.

Context flexing: For all BALANCED and ADVANCED models, the tail of your context window may flex to optimize cache hits. This means that when sending a message after returning to a story, context window may be reduced, but as you continuously write, the context window expands until full (then restarts the process). It flexes the last 30% for most models and 50% for Claude. Cache TTLs:

  • 5 minutes - default
  • 15 minutes - OpenAI, MiMo
  • 1 hour - DeepSeek, Claude Experimental

Transparently, this is a cost reduction measure, but has the side benefit of reducing response latency.

  • Some GPT models are served exclusively via an experimental provider, which has no max output cap at all. Look for the ♾️ badge.

Rate Limits

TierBASIC DailyPRO DailyPRO Monthly
LITE150unlimitedunlimited
BALANCED20unlimitedunlimited
ADVANCED-1502000

Model Descriptions

ModelDescription
Gemma 4 31B TGoogle's Gemma 4 31B, free in Lite via Poe (Google, with Together AI failover).
Ling 2.6 FlashFast, token-efficient model from InclusionAI.
LunarisRoleplay and writing fine-tune of Llama 3 8B
NemoFantastic writer and RPer for its size
Ox AlphaMystery preview model on OpenRouter - LITE for as long as it's available.
DeepSeek v3.2Beloved older DeepSeek, still going strong for RP and writing
DeepSeek V4 FlashLightweight variant of DeepSeek V4
GLM 4.7 FlashFast, lightweight variant of GLM 4.7 from Z.ai
Grok 4 FastEfficient SOTA model by xAI.
Grok 4.1 FastLatest "light" Grok model
Hy3Tencent's Hunyuan 3
MiMo V2 FlashOlder lightweight model from Xiaomi
MiMo V2.5Newest lightweight model from Xiaomi
Step 3.5 FlashFast and smart agentic MoE model from StepFun
DeepSeek V4 ProThe long-awaited newest DeepSeek!
Gemini 2.5 ProPrevious frontier model by Google, thinking
Gemini 3 FlashLatest Gemini Flash model from Google
Gemini 3.1 ProLatest Gemini Pro model from Google
GLM 4.6Really good thinking model from Z.ai
GLM 4.7Older GLM model from Z.ai
GLM 5Previous GLM model from Z.ai
GLM 5.1Previous GLM model from Z.ai
GLM 5.2Previous GLM model from Z.ai
GLM 5.3Newest GLM model from Z.ai
GPT-4.1Frontier model by OpenAI
GPT-5.1Previously the 'Codex' variant, but that's been deprecated - it's a bit more censored but figured this was better than retiring it.
GPT-5.4Older GPT-5 model from OpenAI with optional thinking. Seems very verbose!
GPT-5.5Previous GPT-5 model from OpenAI - available via experimental provider only
GPT-5.6 SolOpenAI's flagship GPT-5.6 model.
Grok 4.6xAI's latest reasoning model.
Haiku 4.5Latest Anthropic Haiku model, their small variant
Kimi K2Thinking version of the latest model from Moonshot - crazy rave reviews, SotA level!
Kimi K2.5Older version of Kimi
Kimi K2.6Previous version for Kimi
Kimi K2.7Previous version of Kimi - this is actually Kimi K2.7 Coding but works great for writing!
Kimi K3Brand new flagship model from Moonshot, huge context window and full multimodal support.
LongCat 2.0Meituan's LongCat 2.0, release version of Owl Alpha
MiMo V2 ProOlder flagship model from Xiaomi, formerly Hunter Alpha
MiMo V2.5 ProNewest flagship model from Xiaomi
MiniMax M2.5Previous frontier model from MiniMax
MiniMax M2.7Previous frontier model from MiniMax
MiniMax M3Latest frontier model from MiniMax
Opus 4.5Anthropic's previous flagship model - available via experimental provider only
Opus 4.6Anthropic's previous flagship model
Opus 4.7Anthropic's previous flagship model. Uses a less token-efficient tokenizer than 4.6 and earlier, so older Claude versions fit a bit more context and output.
Opus 4.8Anthropic's previous flagship model. Uses a less token-efficient tokenizer than 4.6 and earlier, so older Claude versions fit a bit more context and output.
Opus 5Anthropic's newest flagship model. Uses a less token-efficient tokenizer than 4.6 and earlier, so older Claude versions fit a bit more context and output.
Sonnet 4.5Previous Anthropic Sonnet model, highly regarded
Sonnet 4.6Previous Anthropic Sonnet model, highly regarded, thinking-only
Sonnet 5Anthropic's latest Sonnet - near-Opus quality for writing and reasoning at a lower price point. Uses the same less token-efficient tokenizer as Opus 4.7/4.8, so older Claude versions fit a bit more context and output.

I've "jailbroken" these models pretty well, but as I always say, it's a spectrum, and there are no absolutes. For rather extreme content (or just bad luck), even very stable jailbroken models may refuse. The "break glass" extremely uncensored option is probably Deepseek with thinking off. And FYI, don't let a refusal hang around in context - you can retry with a different model and it'll be like it never happened. Or even edit the response and change what the AI said!

Privacy

We gather NO data on you. Your email is used purely for contact and login purposes. Stories are stored purely for multi-device convenience, and we'll introduce 100% client side storage if there's enough demand. Requests are passed to OpenRouter, nano-GPT, or Poe API, which aggregators - please view their privacy policies for further information. Bottom line being it's fully anonymous. Some people have asked if OpenAI can ban you (as we have no choice but to ultimately use their API for OpenAI models), and that's a definite no, they can't, because they have no way of knowing who you are!

Get Started

Ready to spice up your writing? Head to our welcome page to start your first conversation, or jump right into writing.