Practise system design the way the interview tests it.
You can build a job queue, a payment flow, a WebSocket app. The interview asks something else: why it is built that way, what happens when it fails, and what you would change at ten times the load.
Free and without signup. Nothing is graded by AI.
Deploy kills a worker mid-transcode
A routine deploy replaces worker containers. One had been transcoding video 812 for 14 minutes. The container receives SIGTERM and is hard-killed 30 seconds later. What happens to the job now?
Reason through itEvery investigation is one system, built up from its requirements, in the same loop.
- Learn
Short explanations with quick checks: estimate the numbers, predict what breaks, then see why.
- Decide
Choose between designs that all sound plausible, using what you just worked out.
- Break it
A worker dies, a message arrives twice, a request times out after it succeeded.
- Change it
Ten times the traffic, a new requirement, a promise from sales.
- Defend it
Say what the design guarantees, what it assumes, and where it stops working.
After each answer you see why every option is right or wrong, and the few points worth remembering. Choices are checked against a key; if you write down your reasoning, you mark it yourself against specific points. New to the format? Read how the system design interview works.
Systems to design
Solid parts are given; you design the outlined ones.
- A URL shortener like bit.lyDesign a URL ShortenerShort links that never collide, redirect quickly for users anywhere, count every click, and can be switched off in seconds. You will estimate the load first, then design each part from the numbers.Foundational, 45 min9 stages
- A reliable video processing pipelineDesign a Video Processing PipelineInstructors upload multi-gigabyte lectures that take minutes to transcode. Workers crash, deploys interrupt jobs, and the same job can run twice. Every accepted upload must end in exactly one correct, visible outcome.Intermediate, 55 min14 stages
- Rate limiting a public APIDesign an API Rate LimiterA public API must hold every customer to their plan across thirty stateless servers, absorb honest bursts, stop abuse, and never let the limiter itself become the outage.Intermediate, 40 min10 stages
- A home timeline at 300,000 reads a secondDesign a News Feed (Twitter Timeline)Built from Twitter's own account of its home timeline: decide when the work of 'one post, many followers' happens, survive accounts with thirty million followers, and stop precomputing timelines for people who never come back.Intermediate, 45 min9 stages
- A job queue that keeps working when workers fall behindDesign a Distributed Job QueueBuilt from Slack's account of the outage that made it rebuild its job queue: make enqueues safe when workers fall behind, keep one slow job type from starving the rest, and drain a backlog without causing the next outage.Intermediate, 45 min9 stages
- Product analytics over billions of eventsDesign a Product Analytics SystemTake in billions of product events a day and answer questions nobody planned for in seconds: choose where events live, keep ingestion alive through spikes and outages, make filters on people fast, and decide what happens when two anonymous visitors turn out to be one person.Intermediate, 45 min9 stages
One idea from each company
Each system pairs with a company that solved the same problem and wrote about it. Read the original once you have made your own decisions.
All companies- DubRedirects at the edge
- MuxVideo that plays before it is processed
- CloudflareRate limiting at the edge
- Twitter (X)Fan-out for home timelines
- SlackJob queues under backlog
- UberPush instead of polling
- StripePayments that never charge twice
- Meta (Facebook)Caching in front of a database
- DiscordStoring chat history forever
- FigmaReal-time multiplayer editing
- NotionSharding Postgres without downtime
- ConvexA database that pushes query results
- PostHogAnalytics on billions of events
Learn by working it out
Every idea comes in small steps, each with a check that answers straight away. Try the start of the caching lesson.
A cache keeps a copy of data somewhere faster or closer than the original: in the app's memory, in Redis, or on a CDN server near the user. A read checks the cache first. A hit is answered from the copy. A miss goes to the source, and usually stores the answer for next time.
Work it out
A database read takes 10 ms and a cache read takes 1 ms. A miss costs both (check the cache, then read the database). With a 90% hit rate, what's the average read time?Check
Same timings, but each item is read about once, so the hit rate is 5%. What does the cache do?
Then, your own system
Describe a system you built, its parts and how requests flow through it, and get the questions an interviewer would ask about it: trace this request, what survives a restart, what happens when this dependency is slow.