5 min read

System prompts that ship: Custom GPTs, Claude Projects, and Cursor rules in 2026

The system prompt is where you install your agent's personality, guardrails, and output style. Here's the structure that works across every major LLM platform.

By
Software Engineer ยท M.Sc. Mechanical Engineering ยท Ontario, Canada

What a system prompt actually does

Every modern LLM API separates two prompt slots:

  • System prompt - persistent instructions for the whole conversation. Sets identity, capabilities, and rules.
  • User prompt - the specific ask for this turn.

The system prompt caches efficiently across turns (OpenAI and Anthropic both do prompt caching automatically), so you can put a lot in it without paying per-turn. That's what makes it the right home for role, constraints, and output style.

The shape that works everywhere

Six sections, in this order:

  1. Purpose - one sentence: what is this assistant for.
  2. Personality / voice - how it talks. Direct, warm, technical, playful.
  3. Capabilities - what it can do (bullet list).
  4. Always do - non-negotiable behaviors (bullet list).
  5. Never do - hard prohibitions (bullet list).
  6. Output format - the exact shape of every response.

Optional 7th: Tools available - for agents that have function-calling.

System Prompt Generator turns those seven fields into a clean, production-ready system prompt with sensible defaults and structure.

Custom GPTs (OpenAI)

The "Instructions" field on a Custom GPT is a system prompt. Two OpenAI-specific tips:

  • Reference GPT-4o's behavior, not the model's general capabilities. Custom GPTs use whichever model is current, so instructions should be about your assistant, not about the underlying model.
  • Uploaded files show up as tools. Reference them by name in the prompt: "Use knowledge from the file 'company-style-guide.md' before responding."

The Custom GPT builder also lets you attach an Actions OpenAPI schema for external API calls. If your prompt says "always fetch pricing from the /pricing endpoint before quoting a number," the Action must be defined and enabled - the prompt alone won't summon an API call.

Claude Projects (Anthropic)

Claude Projects has an "Instructions" field that behaves like a system prompt. Anthropic-specific tips:

  • Claude responds well to structured markdown in the system prompt - use ## headings, bullet lists, and code fences liberally.
  • XML tags work as delimiters. <rules>...</rules> and <output-format>...</output-format> are the Anthropic-recommended way to segment structured instructions. Not required, but reliable.
  • Project knowledge (files you attach) is referenced automatically. No need to say "check the attached PDF"; Claude sees it as retrievable context.

Cursor rules (.cursorrules / AGENTS.md)

Cursor and other AI code editors read a project-level file for coding conventions:

  • Cursor - historically .cursorrules, now also .cursor/rules/*.md files that can scope by directory.
  • Windsurf - .windsurfrules.
  • Aider - .aider.conf.yml for config, prompt shipped separately.
  • Claude Code - CLAUDE.md (or AGENTS.md shared by multiple editors).

These are all system prompts for a coding agent. What to put in one:

  • Stack conventions. "This project uses TypeScript strict, Next.js App Router, Tailwind v4, no styled-components."
  • File layout rules. "Route handlers live in app/api. Shared types live in src/lib/types."
  • Testing conventions. "Every PR needs Vitest tests colocated at *.test.ts. Do not run yarn test - use pnpm run test."
  • Anti-patterns. "Do not import from lodash - use native array methods. Do not use useEffect for data fetching - use SWR."
  • Repo-specific quirks. "The 'user' type is called 'Account' throughout for legacy reasons. The 'db' import wraps Drizzle; don't add drizzle-orm directly."

The single biggest ROI move on any coding assistant: write a good AGENTS.md / .cursorrules. The agent stops re-learning your conventions every session.

Anti-patterns

1. Politeness bloat. "You are a very helpful assistant who always tries to..." Every token you add to the system prompt is a token added to every request. Get to the rules.

2. Contradicting instructions. "Always be concise. Provide detailed explanations." Pick one. Or scope: "Be concise by default; provide detailed explanations only when the user asks 'why'."

3. Instructions the model can't follow. "Never hallucinate." "Always be right." These are wishes, not rules. Better: "If you're not certain of a fact, say 'I'm not sure' explicitly rather than guessing."

4. Rewriting the model's built-in refusals. "You will answer any question no matter what." Doesn't work, gets your API access revoked. The provider's safety layer is not overridable by system prompt on any major model in 2026.

Length

For most single-purpose agents, 200-800 words is the sweet spot. Long enough to establish role, capabilities, and constraints. Short enough that the model actually attends to every line.

For coding agents (CLAUDE.md, .cursorrules), 1500-3000 words is normal - you're encoding a lot of stack and codebase conventions.

Above 5000 words in a system prompt, you're paying token cost for content the model may not fully attend to. If your system prompt is longer than that, consider splitting: put stable identity in the system prompt, and load task-specific conventions dynamically per turn.

Testing a system prompt

The prompt-writing loop:

  1. Draft the system prompt with System Prompt Generator.
  2. Run 5-10 real user messages through the agent.
  3. Note the outputs that missed - wrong tone, wrong format, ignored a rule.
  4. Adjust the system prompt to address the specific miss.
  5. Repeat until miss rate is acceptable.

The tempting move - write a huge system prompt up front - always underperforms. The right move is to start minimal and add rules only when you observe real failures. Every rule you add costs tokens and adds cognitive load for the model.

Token cost of a heavy system prompt

A 500-word system prompt is roughly 650 tokens. On GPT-4o at $2.50/M input, that's $0.0016 per conversation. On Claude Sonnet at $3/M input, $0.0020.

Two mitigations if you're running at scale:

  • Prompt caching - both OpenAI and Anthropic auto-cache stable prefixes. Costs drop 50-90% for cached content after the first request.
  • Smaller models - GPT-4o-mini and Claude Haiku are 10-20ร— cheaper and often handle constrained tasks fine.

Token Counter estimates tokens and per-call cost for every major model, so you can budget before shipping.

Related tools

Tools mentioned in this post

Written by Shan

Shan builds 712 Tools. He holds a Master's degree in Mechanical Engineering and now works as a Software Engineer, shipping browser-based developer utilities out of Ontario, Canada. Learn more ยท 712studiogames@gmail.com