6 min read

How to write better AI prompts that actually work (ChatGPT, Claude, Gemini in 2026)

Prompt engineering isn't magic - it's a small set of structural moves that produce dramatically better output from any modern LLM. Here's the shape, the anti-patterns, and the template.

By
Software Engineer ยท M.Sc. Mechanical Engineering ยท Ontario, Canada

The single biggest thing separating good prompts from bad

Good prompts specify role, context, task, constraints, and output format. Bad prompts specify only the task.

Compare:

Bad: "Write a landing page for a SaaS."

Good: "You are a landing-page copywriter with a preference for concrete, benefits-first prose. My product is a Chrome extension that saves the text of every article you read to a searchable local database. The audience is knowledge workers who read 20+ articles a week. Write a hero section (headline, subheadline, primary CTA) and three feature blocks. Constraints: no hype adjectives, no 'unleash', use active voice. Format as Markdown."

The bad prompt gets you generic startup copy. The good prompt gets you copy shaped for your product. The structural moves - role, context, task, constraints, format - are what LLMs use to disambiguate the space of possible responses.

The template that works everywhere

Every serious prompt engineer converges on roughly the same shape:

# Role
You are [specific role]. [1-2 lines on approach].

# Context
[The situation, the audience, the constraints of the world.]

# Task
[The actual thing to do - one sentence.]

# Constraints
- [What must always be true]
- [What must never happen]
- [Style / tone rules]

# Output format
[Exactly what the response should look like.]

# Examples (optional)
[1-3 short input โ†’ output pairs.]

That six-section template works on GPT-5, Claude, Gemini, Llama, and every open-source model of consequence. Prompt Enhancer applies it to any rough prompt with one click - pick a preset (coding, writing, analysis, research, image) and it structures your one-liner into the full shape.

The three anti-patterns

1. Politeness inflation. "Could you please help me if it's not too much trouble to..." Every token you add for politeness is a token the model has to attend to. Modern LLMs don't reward manners; get to the point.

2. Vague constraints. "Make it good" is not a constraint. "Under 100 words, no exclamation marks, no adverbs" is a constraint. Constraints that a downstream automated check could verify are constraints worth writing.

3. Buried instructions. Putting critical rules at the end of a long prompt is a common miss - model attention is real, and instructions in the middle of a wall of context get less weight. Front-load anything the model must do; put optional background at the end.

System prompts vs user prompts

Two levers, different jobs:

  • System prompt - persistent instructions that apply to every message in a conversation. Custom GPTs, Claude Projects, Cursor rules, the system parameter of any API. Set the role and constraints here.
  • User prompt - the specific ask for this turn. Task and per-turn context go here.

Common mistake: treating every user message as a fresh full prompt with role + rules. That works but wastes tokens (the system prompt caches efficiently across turns; the user prompt doesn't).

For anything you'll invoke more than twice, extract the stable parts into a system prompt. System Prompt Generator builds those from a form - you fill in agent identity, capabilities, always-do rules, never-do rules, and get a production-ready system prompt for Custom GPTs, Claude Projects, or a .cursorrules file.

Few-shot examples: when they help, when they don't

Including 1-3 example input/output pairs ("few-shot") dramatically improves outputs when:

  • The task has a specific format (JSON structure, headline style, code convention).
  • The judgment is subjective ("write it in this voice").
  • The model tends to default to a boring shape without a nudge.

Few-shot examples don't help - and often hurt - when:

  • The task is straightforward and the model already knows the format.
  • Your examples are inconsistent (mixed signals reduce quality).
  • The examples are so long they crowd out the actual task.

Rule of thumb: two consistent examples beat five inconsistent ones.

Chain-of-thought and its 2026 replacement

"Let's think step by step" (chain-of-thought prompting) was the 2023-2024 way to nudge better reasoning. In 2026, most flagship models - Claude, GPT-5, Gemini - have built-in reasoning modes. Explicitly asking for step-by-step is often redundant and adds token cost. Use it only on smaller models that lack native reasoning (7B and below).

Structured output: JSON schemas, not prompts

Asking "reply as JSON with keys foo and bar" works about 90% of the time. The 10% failure rate - the model wraps output in backticks, adds explanatory prose, omits a field - is exactly why you built a structured pipeline in the first place.

Modern LLMs support structured output via JSON Schema: pass a schema to the API, get guaranteed-valid JSON back. OpenAI, Anthropic, and Google all support this. JSON Schema Generator turns a sample JSON object into a draft-2020-12 schema you can drop into any of those API calls.

Cost matters

Every token you send is billed. Every token the model generates is billed (usually more, since outputs are 2-4ร— the price of inputs).

Prompt engineering that reduces cost:

  • Trim the context. Are you passing an entire 100KB document when the model only needs one paragraph? Extract first.
  • Cache stable prefixes. OpenAI and Anthropic both cache the system prompt and long context across calls. Put your stable content first, dynamic content last.
  • Choose the smaller model. Haiku, GPT-4o-mini, Gemini Flash all cost 10-20ร— less than the flagship - and are often good enough for extraction, classification, and simple rewriting.

Token Counter estimates token count and cost per call for every major model - useful before you build anything that will run at scale.

The two-step pattern that outperforms one-shot

For any hard task, two shorter prompts beat one long one:

  1. Plan - "Given this task, list the sub-tasks you'd do to solve it."
  2. Execute - "Great, now do sub-task 1. [Then 2, then 3.]"

The plan stage lets you catch a misunderstanding cheaply - before the model runs the full expensive execution. This is the pattern behind most agent frameworks (LangGraph, CrewAI, AutoGen).

The workflow

The tools you'll open in the same session:

  1. Prompt Enhancer - turn your rough prompt into the six-section shape.
  2. System Prompt Generator - when the workflow is repeatable, extract stable parts into a system prompt.
  3. Token Counter - sanity-check the token count before you invoke expensive models at scale.
  4. JSON Schema Generator - when you need structured output guarantees.

Good prompting is boring. It's not about clever tricks; it's about clear specification, consistent structure, and measuring what actually shipped.

Tools mentioned in this post

Written by Shan

Shan builds 712 Tools. He holds a Master's degree in Mechanical Engineering and now works as a Software Engineer, shipping browser-based developer utilities out of Ontario, Canada. Learn more ยท 712studiogames@gmail.com