Cost
What running an app and an agent really costs — hosting, cloud dev environments, Claude plans — and how to keep the bill down.
17 posts
Prompt Caching Explained: Cut LLM Costs and Latency on Repeated Prompts
If every request repeats the same long system prompt, documents or tool definitions, prompt caching lets the provider reuse that work for a fraction of the price and time. How it works, structuring prompts for cache hits, Claude's cache_control, OpenAI's automatic caching, and verifying it works.
How to Price Your SaaS: A Beginner's Guide to Charging for Your App
Pricing is the decision first-time founders avoid longest. How to pick a pricing model (flat, tiered, per-seat, usage), how to set a first number, why free plans are expensive, how AI costs change the maths, and how to raise prices later.
Claude Opus vs Sonnet vs Haiku: Which Model Should You Use?
Anthropic's Claude models come in tiers — Opus, Sonnet and Haiku, plus Fable at the top. What each is good at, what they cost on the API, and a simple rule for choosing for coding, chat and features inside your own app.
Claude Code Usage Limits Explained: What "Limit Reached" Means and What to Do
Hit "usage limit reached" in Claude Code? How the five-hour session window and weekly limits work on Pro and Max, how to check what you've used with /usage, and practical ways to make your allowance last.
Claude API Pricing Explained: What Your AI Feature Will Actually Cost
How Claude API billing works: per-token input and output prices for each model, prompt caching, the 50% Batch discount, and worked examples for a chatbot and a summariser — so you can estimate your bill before you ship.
Claude Code /compact vs /clear: Managing Context Without Losing Your Place
When to compact, when to clear, and when to let auto-compaction handle it. What /compact actually does, custom compaction instructions, the auto-compact window and /autocompact, why compacting a huge session is itself expensive, and a workflow that keeps sessions sharp.
How to Change the Model in Claude Code (and Which One to Pick)
Switch models in Claude Code with /model, the --model flag, an environment variable or settings.json. What the aliases mean — opus, sonnet, haiku, fable, opusplan, [1m] — which is the default, and how to choose for cost and quality.
What Is "the Cloud"? Cloud Computing for Beginners
The cloud is other people's computers, rented by the minute. What cloud computing actually is, the difference between IaaS, PaaS, and SaaS, what AWS, Google Cloud, and Azure sell, how cloud billing works, and how much of it a first app really needs.
What Is Rate Limiting? Protecting Your App From Too Many Requests
Rate limiting caps how many requests someone can make in a period. Why every public app needs it — for login forms, sign-ups, AI features, and APIs — how it works, what a 429 response means, where to add limits, and what to do when you hit someone else's.
What Are Tokens in AI? Why Your Usage Is Counted in Pieces of Words
AI models read and write in tokens, not words — and pricing, limits, and context windows are all measured in them. What a token is, roughly how many words it equals, input vs output tokens, why code uses more of them, and practical ways to use fewer.
VPS vs PaaS: Where Should a Small App Live?
A VPS is cheap and does whatever you tell it — including nothing when it breaks. A PaaS runs your app for you and bills you for the privilege, often by usage. What each one actually includes, what it quietly leaves to you, and how to decide for a side project, an AI-built app, or a small business.
How to Stop Bots From Running Up Your AI App's Bill
If your app calls an AI model on a user's behalf, every request costs you money — and a bot, a scraper, or one determined user can make thousands of them overnight. Rate limits, usage caps, provider spending limits, and the architecture that keeps a surprise bill from happening.
How to Keep Claude Code Costs Down (Without Making It Worse)
Claude Code usage is driven less by how much you ask and more by how much context every request carries. Where the tokens actually go, how to see them, and the habits that cut usage — clearing between tasks, picking the right model, trimming CLAUDE.md, and planning before building.
How Much Does It Cost to Run an App? A Realistic Monthly Budget
Hosting, database, domain, email, AI usage, storage, and the tools around them. What each costs for a small app, what's free, where surprise bills come from, and three example budgets — from a side project to a small business with paying customers.
How Much Does GitHub Codespaces Actually Cost?
Codespaces bills by the hour for compute and by the GB-month for storage, with a free allowance on personal accounts. How the meter works, three worked examples from occasional to full-time use, and the settings that stop the bill surprising you.
Claude Pro vs Max vs API Key for Claude Code: Which Should You Pay For?
Claude Code works with a Pro subscription, a Max subscription, or pay-as-you-go API billing. They're metered differently and suit different ways of working. How to choose, with a simple way to check your own usage.
The Real Cost of a Cloud Development Environment
Sticker price is the smallest part of what a cloud development environment costs — and the laptop it replaces was never free either. A model for working out whether the numbers actually favour you.