What you will build
By the end of this tutorial you will have a running GraphQL API for a small bookshelf, served by Apollo Server and written in TypeScript. It has two types, Book and Author, with a real relationship between them, a query that reads books, a query that reads one author by id, a mutation that adds a book and rejects bad input, and a DataLoader fix for the repeated author lookups that the relationship causes.
Data lives in two plain arrays in memory. There is no database to install and nothing to configure, which keeps the focus on the part that is actually new if you are coming from REST: how a schema and a set of resolvers replace a list of routes. The data layer is written behind one function, findAuthorsByIds, that logs every time it is called. That log line is what makes the N+1 section concrete instead of theoretical: you count the lines before the fix and after.
Every file you need is in this page. Paste each one into a fresh project, follow the setup steps, and it runs. There is no repo to clone and no hidden config.
Why GraphQL exists next to REST
GraphQL exists because in REST the server decides the shape of every response, and GraphQL moves that decision to the client. That one change is the whole idea, and almost everything else about GraphQL follows from it.
In a REST API you write a route per resource. GET /books returns whatever fields you decided /books returns. If a mobile screen only needs titles, it still downloads the full book objects. If it also needs each author's name, it either gets a second round of requests (one per book, the classic waterfall) or you add a bespoke GET /books?include=author parameter, and then another one next quarter for the next screen. The API grows a route or a flag per client need.
In GraphQL you publish one endpoint and a typed schema describing everything that can be asked for. The client sends a query naming exactly the fields it wants, including fields on related objects, and gets back a response in that shape. Need titles only? Ask for titles. Need titles plus author names in one trip? Ask for both. The server does not change.
The honest cost is that the work does not disappear, it moves. Letting clients ask for nested fields means the server now resolves those fields on demand, and the naive way to do that fires one lookup per parent object. That is the N+1 problem, and it is the single most common way a GraphQL API gets slow. REST never has to solve it because the server already decided to join the data. GraphQL hands you the flexibility and the bill for it in the same move, which is why the last section of this tutorial is about paying that bill properly rather than an optional extra.
When is GraphQL the right call? When several different clients read the same data in different shapes, and the cost of keeping bespoke endpoints in sync with them is real. When you have one client and a handful of endpoints, a REST API with Express is less machinery for the same result, and that is not a consolation prize.
Set Up the Project
Create a fresh project and install Apollo Server with its graphql peer dependency. Those two are the only runtime packages until the last section. TypeScript, tsx, and @types/node are dev tooling for running and compiling the .ts files. Apollo Server 5 needs Node 20 or newer, so check node --version first.
mkdir bookshelf-graphql
cd bookshelf-graphql
npm init -y
npm install @apollo/server graphql
npm install -D typescript tsx @types/nodeAdd a tsconfig.json at the project root:
{
"compilerOptions": {
"target": "ES2022",
"module": "NodeNext",
"moduleResolution": "NodeNext",
"strict": true,
"esModuleInterop": true,
"skipLibCheck": true,
"rootDir": "src",
"outDir": "dist"
},
"include": ["src"]
}Then open package.json, add "type": "module", and set three scripts. dev runs the server straight from source with tsx and restarts on save, build compiles to dist/, and start runs the compiled output the way a production host would:
{
"name": "bookshelf-graphql",
"private": true,
"type": "module",
"scripts": {
"dev": "tsx watch src/index.ts",
"build": "tsc",
"start": "node dist/index.js"
}
}One detail that trips people up: because this is a real ESM project, every relative import below ends in .js even though the file on disk is .ts. You are importing the path of the compiled output, which is what Node resolves at runtime. tsc and tsx both understand it, and leaving the extension off works under tsx but breaks npm start after a build. Write the .js and it works in both.
Define the GraphQL Schema
The schema is a block of type definitions written in GraphQL's Schema Definition Language, and it is the contract every incoming query is checked against before any of your code runs. Create src/schema.ts:
export const typeDefs = `#graphql
type Author {
id: ID!
name: String!
}
type Book {
id: ID!
title: String!
authorId: ID!
}
type Query {
books: [Book!]!
author(id: ID!): Author
}
`;Reading that top to bottom: Author and Book are object types with scalar fields. ID and String are built-in scalars, and a trailing ! means non-null, so title: String! promises a string on every book and never null. [Book!]! is a non-null list of non-null books, which is the shape you almost always want for a collection: the list itself is always there, and no element inside it is ever null.
Query is special. It is the entry point, the one type whose fields a client is allowed to ask for at the top level of a query. books takes no arguments. author(id: ID!) takes a required id and returns Author with no !, which is a deliberate promise that asking for a missing author is a normal outcome returning null, not an error.
The #graphql comment on the first line is not GraphQL syntax, it is a marker that editors and the GraphQL VS Code extension look for so they syntax-highlight the string. Harmless if your editor ignores it, useful if it does not.
This is what people mean by schema-first design, and it is the sharpest difference from REST in day-to-day work. A REST endpoint's contract is implicit: it lives in a document, or a Swagger file someone has to remember to update, or in the shape of whatever JSON the handler happens to return today. Here the contract is executable. Ask for a field that does not exist and the query is rejected before a resolver is called. Rename a field and every query using the old name fails loudly instead of quietly receiving undefined. The schema is both the documentation and the validator, and it cannot drift from the server because it is the server.
Write the Resolvers
A resolver is a function that returns the value for one field, and the resolver map you pass to Apollo Server mirrors the shape of your schema: an object keyed by type name, then by field name. Start with the data, in src/data.ts:
export interface Author {
id: string;
name: string;
}
export interface BookRecord {
id: string;
title: string;
authorId: string;
}
export const authors: Author[] = [
{ id: "1", name: "Ursula K. Le Guin" },
{ id: "2", name: "Octavia E. Butler" },
];
export const books: BookRecord[] = [
{ id: "1", title: "A Wizard of Earthsea", authorId: "1" },
{ id: "2", title: "The Dispossessed", authorId: "1" },
{ id: "3", title: "Kindred", authorId: "2" },
];
// Stands in for a database query. In a real app this is one round trip:
// select * from authors where id = any($1). The log line is how you count
// those round trips in the N+1 section below.
let lookupCount = 0;
export async function findAuthorsByIds(
ids: readonly string[],
): Promise<Author[]> {
lookupCount += 1;
console.log(
`[data] author lookup #${lookupCount} for ids: ${ids.join(", ")}`,
);
return authors.filter((author) => ids.includes(author.id));
}findAuthorsByIds takes a list of ids rather than a single id on purpose. That signature is what makes batching possible later, and it is the one design decision in this file worth copying into real code: write your data layer to accept a set of keys, and the N+1 fix becomes a change in the resolver instead of a rewrite of the data access.
Now the schema needs one more field. Right now Book exposes a raw authorId, which pushes the work of turning that id into a name back onto the client. Add an author field to the Book type in src/schema.ts:
type Book {
id: ID!
title: String!
authorId: ID!
author: Author!
}That field has no column behind it. Nothing in BookRecord is called author. It exists only because a resolver will produce it, which is the thing to notice about a GraphQL schema: it describes a graph your clients can walk, not the tables you happen to store. Create src/resolvers.ts:
import { GraphQLError } from "graphql";
import { books, findAuthorsByIds, type BookRecord } from "./data.js";
export const resolvers = {
Query: {
books: () => books,
author: async (_parent: unknown, args: { id: string }) => {
const [author] = await findAuthorsByIds([args.id]);
return author ?? null;
},
},
Book: {
author: async (book: BookRecord) => {
const [author] = await findAuthorsByIds([book.authorId]);
if (!author) {
throw new GraphQLError(`No author with id ${book.authorId}`);
}
return author;
},
},
};Three resolvers, and they are doing three different jobs.
Query.books returns the array. Every scalar field on those objects resolves itself: GraphQL's default resolver reads the matching property off the parent, so you never write a resolver for Book.title. You only write one when the field needs work.
Query.author shows the argument signature. A resolver's second parameter is the field's arguments, typed here as { id: string }. Returning null for a missing author is legal because the schema declared author: Author as nullable, and it is the right answer: a client asking about an id that is not there has not triggered an error, it has learned something.
Book.author is the interesting one. Its first parameter is the parent object, the specific BookRecord currently being resolved, and that is how a nested field finds its context. GraphQL walks the query depth-first: resolve books, then for each book in the result, resolve author. One call per book. Hold that thought, because it is the N+1 problem in one sentence and the last section is about it.
Throwing GraphQLError is the right way to fail. Apollo catches it and returns a structured errors array alongside whatever data did resolve, which is another REST difference: there is no single status code for the whole response, because half a GraphQL response succeeding is a normal outcome.
Start the Apollo Server
startStandaloneServer boots an HTTP server with your schema mounted on a single endpoint and the Apollo Sandbox explorer attached, which is all you need in development. Create src/index.ts:
import { ApolloServer } from "@apollo/server";
import { startStandaloneServer } from "@apollo/server/standalone";
import { typeDefs } from "./schema.js";
import { resolvers } from "./resolvers.js";
const server = new ApolloServer({ typeDefs, resolvers });
const { url } = await startStandaloneServer(server, {
listen: { port: 4000 },
});
console.log(`Bookshelf API ready at ${url}`);Run it:
npm run devYou should see Bookshelf API ready at http://localhost:4000/. That one URL is the entire API surface. There is no /books, no /authors, no versioned prefix. Every read and every write goes to the same path as a POST, and the query body says what you want.
The top-level await works without any wrapper because this is an ES module, which is why "type": "module" was in package.json. startStandaloneServer is the quick path and it is deliberately limited: no custom middleware, no extra routes, no control over CORS beyond the defaults. When you need any of that, you swap it for expressMiddleware and keep every other file in this tutorial unchanged. The FAQ at the bottom has the shape of that swap.
Query the API
Open http://localhost:4000/ in a browser and Apollo Sandbox loads, an explorer that reads your schema and autocompletes every field as you type. Paste this into the left pane and run it:
query {
books {
title
author {
name
}
}
}The same query over curl, which is what your client code will actually do:
curl -s http://localhost:4000/ \
-H 'Content-Type: application/json' \
-d '{"query":"{ books { title author { name } } }"}'{
"data": {
"books": [
{ "title": "A Wizard of Earthsea", "author": { "name": "Ursula K. Le Guin" } },
{ "title": "The Dispossessed", "author": { "name": "Ursula K. Le Guin" } },
{ "title": "Kindred", "author": { "name": "Octavia E. Butler" } }
]
}
}Look at what came back and, more to the point, what did not. No id, no authorId, no author id. The records in data.ts carry those fields and the schema exposes them, but the query did not ask, so they were not sent. Drop author from the query and you get three titles and nothing else. Add id and you get ids. The response shape is the query shape, every time.
That is the client-driven field selection REST cannot give you without a parameter for every combination. A mobile list view asks for titles. The detail screen asks for titles, ids, and author names. A sitemap job asks for ids alone. Three clients, three payload sizes, one endpoint, and no server change for any of them.
Now look at your server terminal. Three author lookups for three books:
[data] author lookup #1 for ids: 1
[data] author lookup #2 for ids: 1
[data] author lookup #3 for ids: 2Three lookups, and two of them asked for the same author. That is the bill GraphQL hands you for the flexibility above, and the last section pays it.
Go deeper on backend APIs
Want the full backend picture around this? Full Stack Open (University of Helsinki) has a free, project-based Node.js track including a GraphQL section, and Back End Development and APIs (freeCodeCamp) is a good parallel if you prefer a certification path. Both sit on our backend developer roadmap.
Add a Mutation
A mutation is a write, and you declare it on the Mutation type exactly the way a read is declared on Query. Add this to src/schema.ts:
type Mutation {
addBook(title: String!, authorId: ID!): Book!
}Returning Book! rather than a boolean or a bare id matters more than it looks. Because the mutation returns a type, the client picks the fields it wants back from the thing it just created, in the same round trip, with no follow-up read. Then add the resolver to src/resolvers.ts, alongside Query and Book:
Mutation: {
addBook: async (
_parent: unknown,
args: { title: string; authorId: string },
) => {
const title = args.title.trim();
if (title.length === 0) {
throw new GraphQLError("title cannot be empty", {
extensions: { code: "BAD_USER_INPUT", argumentName: "title" },
});
}
const [author] = await findAuthorsByIds([args.authorId]);
if (!author) {
throw new GraphQLError(`No author with id ${args.authorId}`, {
extensions: { code: "BAD_USER_INPUT", argumentName: "authorId" },
});
}
const book: BookRecord = {
id: String(books.length + 1),
title,
authorId: args.authorId,
};
books.push(book);
return book;
},
},The schema already did some of the validation for you. title: String! means a query omitting the title, or sending null, or sending a number, is rejected before this function runs. What the type system cannot know is that a title of three spaces is not a title, so that check is yours. This is the split worth internalizing: types catch shape, resolvers catch meaning.
extensions.code is how you give a client something to branch on. BAD_USER_INPUT is Apollo's conventional code for a rejected argument, and it plays the role an HTTP 400 plays in REST, except it arrives per field rather than per response. Write the message for a human and the code for a program.
Try it:
curl -s http://localhost:4000/ \
-H 'Content-Type: application/json' \
-d '{"query":"mutation { addBook(title: \"Parable of the Sower\", authorId: \"2\") { id title author { name } } }"}'That returns the new book with its author name resolved through the same nested resolver a read uses. Now send an empty title and watch it fail cleanly:
curl -s http://localhost:4000/ \
-H 'Content-Type: application/json' \
-d '{"query":"mutation { addBook(title: \" \", authorId: \"2\") { id } }"}'You get "data": null and an errors array carrying the message and "code": "BAD_USER_INPUT". The HTTP status is still 200, which surprises everyone coming from REST and is the intended design: transport succeeded, the operation did not, and those are separate facts. Clients check errors, not the status line.
Fix the N+1 Problem with DataLoader
The N+1 problem is what you already saw in the server log: one query for the list of books, then one more lookup per book to resolve its author, so N books cost N+1 trips to the data layer. DataLoader fixes it by collecting every id requested during one tick of the event loop and asking for all of them in a single batch.
The log from that three-book query is the whole problem in three lines. Three books, three author lookups, and ids 1 and 1 and 2, so one of those lookups was pure duplication. Here the data is an array in memory and the cost is nothing, which is exactly why this is the right place to learn it: swap findAuthorsByIds for a real select and those are three network round trips to Postgres. Put thirty books on the page and it is thirty-one queries for one screen. This is the most common reason a GraphQL API that looked fine in development falls over in production, and it does not announce itself, because every individual resolver is fast and correct.
It is also structural rather than a mistake someone made. Book.author cannot know it is being called in a loop. It gets one book and returns one author, which is the only honest thing a field resolver can do. The fix has to come from outside the resolver, and that is what DataLoader is.
Install it:
npm install dataloaderCreate src/loaders.ts:
import DataLoader from "dataloader";
import { findAuthorsByIds, type Author } from "./data.js";
export interface Loaders {
author: DataLoader<string, Author | null>;
}
export interface GraphQLContext {
loaders: Loaders;
}
export function createLoaders(): Loaders {
return {
author: new DataLoader<string, Author | null>(async (ids) => {
const rows = await findAuthorsByIds(ids);
const byId = new Map(rows.map((author) => [author.id, author] as const));
// Must return one entry per requested id, in the same order.
return ids.map((id) => byId.get(id) ?? null);
}),
};
}The batch function is the entire contract and it has one rule that causes nearly every DataLoader bug: it receives an array of keys and must return an array of the same length, in the same order, with one entry per key. findAuthorsByIds uses a filter, so it drops missing ids and returns a shorter array in whatever order the source had. Hand that straight back and books silently get each other's authors. Building the Map and then mapping over ids is not defensive style, it is the requirement.
Now wire a fresh set of loaders into every request. Apollo builds a context object per request and passes it to every resolver, which is the right lifetime for a loader: it needs to be shared across all the resolvers in one query, and it must not be shared across two. DataLoader caches each key it has seen, so a long-lived loader would keep serving an author's old name after someone renamed it. Update src/index.ts:
import { ApolloServer } from "@apollo/server";
import { startStandaloneServer } from "@apollo/server/standalone";
import { typeDefs } from "./schema.js";
import { resolvers } from "./resolvers.js";
import { createLoaders, type GraphQLContext } from "./loaders.js";
const server = new ApolloServer<GraphQLContext>({ typeDefs, resolvers });
const { url } = await startStandaloneServer(server, {
context: async () => ({ loaders: createLoaders() }),
listen: { port: 4000 },
});
console.log(`Bookshelf API ready at ${url}`);The <GraphQLContext> on ApolloServer is what makes the context typed inside your resolvers, so context.loaders.author autocompletes and a typo is a compile error rather than a runtime undefined. Finally, point the resolvers at the loader instead of calling the data layer directly. In src/resolvers.ts, swap the imports and both author resolvers:
import { GraphQLError } from "graphql";
import { books, findAuthorsByIds, type BookRecord } from "./data.js";
import type { GraphQLContext } from "./loaders.js";
export const resolvers = {
Query: {
books: () => books,
author: (
_parent: unknown,
args: { id: string },
context: GraphQLContext,
) => context.loaders.author.load(args.id),
},
Book: {
author: async (
book: BookRecord,
_args: unknown,
context: GraphQLContext,
) => {
const author = await context.loaders.author.load(book.authorId);
if (!author) {
throw new GraphQLError(`No author with id ${book.authorId}`);
}
return author;
},
},
Mutation: {
// unchanged from the previous section
},
};Keep the Mutation block exactly as you wrote it, and keep the findAuthorsByIds import, since addBook still uses it to check that the author exists. Restart with npm run dev and run the three-book query again. The log is now one line:
[data] author lookup #1 for ids: 1, 2One lookup, two ids, three books served. Two things happened there. The three .load() calls queued up instead of firing, because DataLoader waits until the current tick of the event loop drains and then calls your batch function once with everything it collected. And the duplicate dropped out: books 1 and 2 share author 1, so the batch asked for 1 and 2, not 1, 1, 2. Batching and per-request caching are separate features and you got both from one change.
The resolvers did not learn anything about each other to make this work. Book.author still takes one book and returns one author. That is the part worth carrying to your own code: the fix lives in a layer the resolvers borrow, so it keeps working as the schema grows, and adding a Author.books field tomorrow needs a second loader, not a redesign.
The habit to build is boring and effective. Any time you write a resolver on a nested field that reaches for data by id, give it a loader on day one, before anything is slow. Retrofitting loaders onto a schema with forty resolvers is a long afternoon. Writing one when you add the field costs five minutes.
Where to take this next
You now have the four pieces every GraphQL server is built from: a schema, resolvers, a per-request context, and batched data access. Swapping the in-memory arrays for a real database changes one file, data.ts, and nothing else, which is the payoff for keeping the data layer behind findAuthorsByIds from the start.
Three directions from here, roughly in the order they pay off:
- Put a database behind it. Rewrite
findAuthorsByIdsas a singlewhere id in (...)query and the DataLoader work you just did starts earning real round trips. Our REST API with Prisma and PostgreSQL tutorial sets up exactly that persistence layer, and the Prisma client slots straight into this file. - Generate your resolver types from the schema. The resolvers here are hand-typed, which is fine at this size and gets tedious fast.
@graphql-codegen/clireads your SDL and emits exact types for every parent, argument, and return value, so a schema change breaks the build in the right place. - Limit query depth and complexity. Client-driven queries mean a client can ask for something enormous, including a cycle through nested fields. Public GraphQL endpoints need a depth or cost limit the same way a REST API needs rate limiting.
If the schema side is what interested you most, that is a signal worth following: schema design is where GraphQL projects are won or lost, well before performance becomes the question. And if you want the contrast fresh in your head, build the same bookshelf as a REST API with Express and TypeScript and compare how each one handles a client that needs one new field.