Skip to Content
ExamplesFrameworksNext.jsCalling Claude

Calling Claude on Next.js

Two endpoints backed by the Claude API that remember the conversation on the server, and a React component that renders each reply as it arrives

Overview

A resource calls Claude the way it does anything else: an ordinary await inside POST. The part that needs thought is the shape of the answer. A model generates its reply a piece at a time, and a request that waits for the whole thing can sit for a minute before anything reaches the browser. So this example builds two endpoints, one that waits and one that streams, plus a component with a tab for each, so you can send the same conversation to either and watch the difference:

EndpointAnswers with
POST /api/chatOne JSON body, once the model has finished
POST /api/chat/streamA server-sent event  per piece, as the model produces it

Both remember the conversation. The browser sends only the new prompt, and the server adds it to the history it kept from the earlier turns. A third endpoint, /api/conversation, returns that history on GET, so a reload picks up where it left off, and forgets it on DELETE, when you start over.

A Next.js route handler — a file exporting one function per HTTP method — is given a Web Request and returns a Response. That is exactly what the chain accepts and returns, so a streamed body needs nothing from Drash: the resource returns a Response whose body happens to be a ReadableStream, and Drash hands it back untouched.

Objectives

To gain familiarity with:

  • mounting a Drash application inside a Next.js route handler;
  • returning a Response whose body is a ReadableStream;
  • turning an API client’s errors into ones the chain can render;
  • reassembling server-sent events in the browser, where a network read can end in the middle of one;
  • remembering a conversation on the server, since the API keeps no state between calls; and
  • letting Claude browse the web with server tools, and showing the sources it cites.

Steps

Folder Structure End State

          • route.ts
      • chat-client.tsx
      • models.ts
      • page.tsx

Set Up the Project

You need a Next.js application — create-next-app makes one — and an Anthropic API key .

npx create-next-app@latest my-app --ts --app --tailwind cd my-app npm install @drashland/drash @anthropic-ai/sdk react-markdown remark-gfm npm install --save-dev @tailwindcss/typography export ANTHROPIC_API_KEY="sk-ant-..."

react-markdown and remark-gfm render Claude’s replies, which are Markdown. @tailwindcss/typography styles what they produce; turn it on by adding one line to app/globals.css, below the @import that create-next-app added:

@import "tailwindcss"; @plugin "@tailwindcss/typography";

This page was written against Next.js 16.3.6, React 19, and Tailwind CSS 4, which is what create-next-app@latest produces today (September 23, 2026). The --tailwind flag installs and configures Tailwind. Only the styling below depends on it, so if you have your own styling, drop the flag, the typography plugin, and the classes.

Keep the key on the server. The route handler and the page shell run on the server. The component marked "use client" runs in the browser and never sees the key — it talks to your own endpoint, which talks to Anthropic. An Anthropic key in client-side code is readable by anyone who opens devtools, and every request they make with it is charged to your account.

Create the Following Files

app/api/[...drash]/route.ts
import Anthropic from "@anthropic-ai/sdk"; import { Application, HTTPError, Resource, } from "@drashland/drash/modules/http.polyfill.js"; import { Status } from "@drashland/drash/core/http/response/Status.js"; import { DEFAULT_MODEL, isModel, type Model } from "@/app/models"; const client = new Anthropic({ apiKey: process.env.ANTHROPIC_API_KEY, }); type Turn = { role: "user" | "assistant"; content: string }; // The API keeps no state between calls, so something has to hold the // conversation and send all of it every time. Here that is the server: one // list of turns per conversation, in memory. A restart forgets them all, and // serverless instances do not share memory, so a real deployment keeps these // in a database or a key-value store instead. const conversations = new Map<string, Turn[]>(); const COOKIE = "conversation"; // The browser holds only an ID, in a cookie, and sends it back on its own. // `HttpOnly` keeps page scripts from reading it. function openConversation(request: Request) { const cookies = request.headers.get("cookie") ?? ""; const found = cookies.match(new RegExp(`(?:^|;\\s*)${COOKIE}=([^;]+)`))?.[1]; const id = found ?? crypto.randomUUID(); const headers: Record<string, string> = found // Only a new conversation ? {} // needs the cookie set. : { "set-cookie": `${COOKIE}=${id}; Path=/; HttpOnly; SameSite=Lax` }; return { id, turns: conversations.get(id) ?? [], headers }; } // Stores the prompt and the reply together, and only once the reply is done. // A failed request leaves no trace, so the next one never sends Claude a // question that was never answered. function remember(id: string, turns: Turn[], prompt: string, reply: string) { if (reply === "") { // The API rejects an empty return; // turn, so storing one would } // break every later request. conversations.set(id, [ ...turns, { role: "user", content: prompt }, { role: "assistant", content: reply }, ]); } // Both chat resources start here. A route handler is handed a Web `Request`, // which parses its own body. async function readRequest( request: Request, ): Promise<{ prompt: string; model: Model }> { let body: { prompt?: unknown; model?: unknown }; try { body = await request.json(); } catch { throw new HTTPError(Status.BadRequest, "Body must be JSON"); } if (typeof body.prompt !== "string" || body.prompt === "") { throw new HTTPError(Status.BadRequest, "Body must have a `prompt` string"); } const model = body.model ?? DEFAULT_MODEL; // The browser picks from a // list, but nothing stops a if (!isModel(model)) { // caller sending any string. throw new HTTPError(Status.BadRequest, "Unknown model"); } return { prompt: body.prompt, model }; } // Anthropic's failures are not your caller's failures. Restate each one as the // status your caller should actually see, with a reason a person can act on. // Order matters — the specific classes come before `APIError`, which is their // shared parent. function anthropicErrorToHTTPError(error: unknown): unknown { if (error instanceof Anthropic.AuthenticationError) { return new HTTPError( // Your key is wrong. The Status.InternalServerError, // caller cannot fix that. "The server's Anthropic API key was rejected.", ); } if (error instanceof Anthropic.RateLimitError) { return new HTTPError( // Pass the limit on. Status.TooManyRequests, "Too many requests right now. Wait a moment, then try again.", ); } if (error instanceof Anthropic.NotFoundError) { return new HTTPError( // A model this key has no Status.BadRequest, // access to. "That model is not available to this server's API key.", ); } if (error instanceof Anthropic.BadRequestError) { return new HTTPError( // Usually the prompt. Status.BadRequest, "Claude could not accept that request. The conversation may be too long.", ); } if (error instanceof Anthropic.APIError) { return new HTTPError( // Upstream is down. Status.BadGateway, "Claude could not be reached. Try again in a moment.", ); } return error; // Not ours. Leave it. } // Claude can search the web and read pages, but only with tools it is given. // Both run on Anthropic's servers: the reply comes back with the browsing // already done, so nothing on this server fetches anything. Haiku 4.5 takes the // older versions. The newer ones let Claude filter results with code before // reading them, which spends fewer tokens. `max_uses` caps the cost per reply. function webTools(model: Model): Anthropic.Messages.ToolUnion[] { if (model === "claude-haiku-4-5") { return [ { type: "web_search_20250305", name: "web_search", max_uses: 5 }, { type: "web_fetch_20250910", name: "web_fetch", max_uses: 5 }, ]; } return [ { type: "web_search_20260209", name: "web_search", max_uses: 5 }, { type: "web_fetch_20260209", name: "web_fetch", max_uses: 5 }, ]; } // Browsing runs as a loop on Anthropic's side, and a long one stops partway // with `stop_reason: "pause_turn"`. Sending the reply so far back picks it up // where it stopped. This caps how many times that happens for one prompt. const MAX_CONTINUATIONS = 3; // Search results come back as citations on the text that used them. Anything // that shows web search results has to show where they came from, so they are // listed under the reply. Brackets would end the link text early. type Sources = Map<string, string>; // URL → page title. function collect(sources: Sources, citation: Anthropic.TextCitation) { if (citation.type === "web_search_result_location") { sources.set(citation.url, citation.title ?? citation.url); } } function listSources(sources: Sources): string { if (sources.size === 0) { return ""; } const links = [...sources].map( ([url, title]) => `- [${title.replace(/[[\]]/g, "\\$&")}](${url})`, ); return `\n\n**Sources**\n\n${links.join("\n")}`; } // A reply is several text blocks when Claude writes, searches, then writes // again. Joined as-is, "Let me check." runs straight into the answer, so a // paragraph break goes wherever text resumes after something that was not // text. Adjacent text blocks are one sentence split by citations; they join // with nothing. const resumes = (previous: string, reply: string) => previous !== "text" && reply !== "" ? "\n\n" : ""; // A refusal is not an exception. The API answers `200` and says so in // `stop_reason`, so it is checked for and turned into one. const refused = () => new HTTPError( Status.UnprocessableEntity, "Claude declined to answer that prompt.", ); class Conversation extends Resource { // The page reads the thread public paths = ["/api/conversation"]; // back from here on load. public GET(request: Request) { const { turns } = openConversation(request); return Response.json({ messages: turns }); } public DELETE(request: Request) { // Starts over. The cookie const { id } = openConversation(request); // stays; its history goes. conversations.delete(id); return new Response(null, { status: 204 }); } } class Chat extends Resource { // Answers with the whole reply. public paths = ["/api/chat"]; // The URL the browser asks // for, not a path relative public async POST(request: Request) { // to the handler. const { prompt, model } = await readRequest(request); const { id, turns, headers } = openConversation(request); const messages: Anthropic.MessageParam[] = [ ...turns, { role: "user", content: prompt }, ]; const sources: Sources = new Map(); let text = ""; let previous = ""; for (let round = 0; ; round++) { let message; try { message = await client.messages.create({ model, max_tokens: 16000, tools: webTools(model), messages, }); } catch (error) { throw anthropicErrorToHTTPError(error); } if (message.stop_reason === "refusal") { throw refused(); } for (const block of message.content) { // `content` holds blocks of if (block.type === "text") { // several types: searches, text += resumes(previous, text) + block.text;// their results, and text. block.citations?.forEach((citation) => collect(sources, citation)); } previous = block.type; } if (message.stop_reason !== "pause_turn" || round === MAX_CONTINUATIONS) { break; } messages.push({ role: "assistant", content: message.content }); } text += listSources(sources); remember(id, turns, prompt, text); return Response.json({ text }, { headers }); } } class ChatStream extends Resource { // Answers a piece at a time. public paths = ["/api/chat/stream"]; public async POST(request: Request) { // Read and checked before the `Response` exists — that matters. const { prompt, model } = await readRequest(request); const { id, turns, headers } = openConversation(request); const messages: Anthropic.MessageParam[] = [ ...turns, { role: "user", content: prompt }, ]; const encoder = new TextEncoder(); const body = new ReadableStream({ async start(controller) { const send = (frame: string) => controller.enqueue(encoder.encode(frame)); let reply = ""; // Collected as it goes out, // so it can be stored once const write = (text: string) => { // the stream is done. if (text === "") { return; } reply += text; send(`data: ${JSON.stringify({ text })}\n\n`); }; const sources: Sources = new Map(); let previous = ""; try { for (let round = 0; ; round++) { const stream = client.messages.stream({ model, max_tokens: 64000, tools: webTools(model), messages, }); for await (const event of stream) { if (event.type === "content_block_start") { const block = event.content_block; if (block.type === "text") { write(resumes(previous, reply)); } if (block.type === "server_tool_use") {// A search can take several const text = block.name === "web_fetch" // seconds with no text, so ? "Reading a page" // say what is happening. : "Searching the web"; send(`event: activity\ndata: ${JSON.stringify({ text })}\n\n`); } previous = block.type; } if (event.type === "content_block_delta") { if (event.delta.type === "text_delta") { write(event.delta.text); } if (event.delta.type === "citations_delta") { collect(sources, event.delta.citation); } } if ( event.type === "message_delta" && event.delta.stop_reason === "refusal" ) { throw refused(); // Any text already sent is } // dropped by the page. } const message = await stream.finalMessage(); if ( message.stop_reason !== "pause_turn" || round === MAX_CONTINUATIONS ) { break; } messages.push({ role: "assistant", content: message.content }); } write(listSources(sources)); remember(id, turns, prompt, reply); send("event: done\ndata: {}\n\n"); } catch (error) { // Too late for a status code. const failure = anthropicErrorToHTTPError(error); const message = failure instanceof HTTPError // Say so in the stream ? failure.message // instead, with the same : "The reply stopped partway through."; // reason /api/chat gives. send(`event: error\ndata: ${JSON.stringify({ message })}\n\n`); console.error(error); } finally { controller.close(); } }, }); return new Response(body, { headers: { ...headers, "content-type": "text/event-stream", "cache-control": "no-store", "x-accel-buffering": "no", // Stops a reverse proxy from }, // holding frames back. }); } } const app = Application .builder() .resources(Conversation, Chat, ChatStream) .build(); function handle(request: Request): Promise<Response> { return app .handle<Response>(request) .catch((error) => { // `name` is checked before `instanceof` because `instanceof` fails when // two copies of Drash end up in one bundle. if (error.name === "HTTPError") { return new Response(error.message, { status: error.status_code, statusText: error.status_code_description, }); } console.error(error); return new Response("The server could not generate a response", { status: 500, }); }); } // Next.js dispatches on the export name, so every method the chain should see // needs one. All of them hand the request straight to the same application. export async function GET(request: Request) { return handle(request); } export async function POST(request: Request) { return handle(request); } export async function DELETE(request: Request) { return handle(request); }
app/models.ts
// The models the page offers and the only ones the server will call. Both // sides import this list, so they cannot disagree. It is plain data with no // secrets in it, which is what makes it safe to ship to the browser. export const MODELS = { "claude-opus-5": "Claude Opus 5", "claude-sonnet-5": "Claude Sonnet 5", "claude-haiku-4-5": "Claude Haiku 4.5", "claude-fable-5-1": "Claude Fable 5.1", } as const; export type Model = keyof typeof MODELS; export const DEFAULT_MODEL: Model = "claude-opus-5"; export function isModel(value: unknown): value is Model { return typeof value === "string" && Object.hasOwn(MODELS, value); }
app/page.tsx
import { ChatClient } from "./chat-client"; // A server component. It renders on the server and ships no JavaScript of its // own, so nothing here is reachable from the browser. export default function ChatPage() { return ( <div className="flex h-dvh flex-col items-center bg-zinc-50 p-6 font-sans dark:bg-zinc-950"> <main className="flex min-h-0 w-full flex-1 flex-col items-center gap-3"> <h1 className="text-3xl font-semibold leading-10 tracking-tight text-black dark:text-zinc-50"> Chat with Claude </h1> <ChatClient /> </main> </div> ); }
app/chat-client.tsx
"use client"; import { useEffect, useRef, useState } from "react"; import Markdown from "react-markdown"; import remarkGfm from "remark-gfm"; import { DEFAULT_MODEL, MODELS, type Model } from "./models"; type Turn = { role: "user" | "assistant"; content: string; error?: string; // "Not delivered", and why. }; // The server never stored it. // One tab per resource. Both add to the same conversation on the server, so // switching tabs partway through carries it over. const ENDPOINTS = { "/api/chat/stream": { summary: "Receive text as the model writes it", explanation: ( <> <p> The server starts its response straight away and forwards each piece of text the moment Claude writes it, as a{" "} <em>server-sent event</em>: a <code>data:</code> line, then a blank line. </p> <pre className="overflow-x-auto rounded-lg bg-black/[.05] p-3 text-xs dark:bg-white/[.08]"> {`data: {"text":"Hel"}\n\ndata: {"text":"lo"}\n\nevent: done`} </pre> <p> The first words show up within a second or two and the reply grows as you watch. The page reads the body piece by piece and stitches the events back together, because a network read can stop in the middle of one. </p> <p> The trade-off: the <code>200</code> status goes out before Claude has written anything, so a failure partway through cannot change it. The server reports it inside the stream instead, as an{" "} <code>event: error</code>. </p> </> ), }, "/api/chat": { summary: "Receive all text once the model has finished", explanation: ( <> <p> The server asks Claude for the whole reply and waits. Only when the model has written its last word does the server answer, with one JSON body: <code>{`{"text": "..."}`}</code>. </p> <p> That makes the client simple: one <code>fetch</code>, one{" "} <code>response.json()</code>. The cost is the wait. Nothing appears until the reply is finished, and a long reply can take a minute. </p> <p> Failures are simple too. Nothing has been sent when something goes wrong, so the status code (400, 429, 500, 502) says exactly what happened. </p> </> ), }, } as const; type Endpoint = keyof typeof ENDPOINTS; export function ChatClient() { const [endpoint, setEndpoint] = useState<Endpoint>("/api/chat/stream"); const [model, setModel] = useState<Model>(DEFAULT_MODEL); const [messages, setMessages] = useState<Turn[]>([]); const [busy, setBusy] = useState(false); const [loaded, setLoaded] = useState(false); const [problem, setProblem] = useState(""); const [activity, setActivity] = useState(""); const thread = useRef<HTMLDivElement>(null); const input = useRef<HTMLInputElement>(null); const learnMore = useRef<HTMLDialogElement>(null); const errorModal = useRef<HTMLDialogElement>(null); const startOverModal = useRef<HTMLDialogElement>(null); useEffect(() => { // The server holds the fetch("/api/conversation") // conversation, so a reload .then((response) => response.json()) // picks up where it left off. .then((body: { messages: Turn[] }) => setMessages(body.messages)) .catch(() => {}) // Nothing to restore; start .finally(() => setLoaded(true)); // empty. }, []); useEffect(() => { // Follow the newest bubble, const box = thread.current; // the way a messaging app box?.scrollTo({ top: box.scrollHeight }); // does. Only the thread }, [messages]); // scrolls, not the page. const locked = busy || !loaded; // No sending until the saved // thread arrives, or it would // overwrite the new prompt. function send(event: React.FormEvent<HTMLFormElement>) { event.preventDefault(); const form = event.currentTarget; const prompt = String(new FormData(form).get("prompt")).trim(); if (!prompt || locked) { return; } form.reset(); submit(prompt); } function retry(index: number) { // The server never stored the const prompt = messages[index].content; // failed prompt, so sending it // again is just sending it. setMessages((turns) => turns.filter((_, i) => i !== index)); input.current?.focus(); // The button is about to submit(prompt); // disappear from under focus. } async function submit(prompt: string) { setMessages((turns) => [ ...turns, { role: "user", content: prompt }, { role: "assistant", content: "" }, // Where the reply lands. ]); setBusy(true); const fail = (reason: string) => { setMessages((turns) => [ // Drop the half-written ...turns.slice(0, -2), // reply and mark the prompt, { ...turns[turns.length - 2], error: reason }, // like iMessage does. ]); setBusy(false); setActivity(""); setProblem(reason); errorModal.current?.showModal(); }; const append = (text: string) => { setMessages((turns) => { // Add to the last bubble. const last = turns[turns.length - 1]; // The function form sees every // earlier append, even ones return [ // that landed between renders. ...turns.slice(0, -1), { ...last, content: last.content + text }, ]; }); }; // Only the new prompt goes up. The server already has the rest, found by // the cookie the browser sends along with the request on its own. const response = await fetch(endpoint, { method: "POST", // `EventSource` only sends headers: { "content-type": "application/json" }, // GET, so the stream is read body: JSON.stringify({ prompt, model }), // by hand below instead. }).catch(() => null); if (!response) { fail("Could not reach the server. Check your connection and try again."); return; } if (!response.ok || !response.body) { // A failure before the reply const reason = await response.text(); // started still has a status fail(reason || `Error ${response.status}`); // code worth believing. return; } if (endpoint === "/api/chat") { // The whole reply, in one go. const { text } = await response.json(); append(text); setBusy(false); return; } const reader = response.body.getReader(); const decoder = new TextDecoder(); let buffer = ""; while (true) { const { done, value } = await reader.read(); if (done) { break; } buffer += decoder.decode(value, { stream: true }); const frames = buffer.split("\n\n"); // A read can end mid-frame, buffer = frames.pop() ?? ""; // so hold the tail back until // its blank line arrives. for (const frame of frames) { const name = frame.match(/^event: (.*)$/m)?.[1] ?? "message"; const data = frame.match(/^data: (.*)$/m)?.[1]; if (!data) { continue; } if (name === "error") { // The stream already sent 200 fail(JSON.parse(data).message); // and then failed. Without return; // this, a failure looks like } // a short, successful answer. if (name === "activity") { // Claude is searching or setActivity(JSON.parse(data).text); // reading, not writing. Say continue; // so under the bubble. } if (name === "done") { setBusy(false); setActivity(""); return; } setActivity(""); // Writing again. append(JSON.parse(data).text); } } // The body ended without a `done` event, so the reply is not finished. fail("The connection closed before the reply finished."); } function confirmStartOver() { const dialog = startOverModal.current; if (!dialog) { return; } dialog.returnValue = ""; // It keeps the last answer, dialog.showModal(); // so a "yes" from before would } // count again on Escape. // Forgotten on the server first, then on screen. async function startOver() { await fetch("/api/conversation", { method: "DELETE" }); setMessages([]); input.current?.focus(); } return ( <div className="flex min-h-0 w-full flex-1 flex-col items-center gap-3"> <div role="tablist" className="flex gap-1 rounded-lg bg-black/[.05] p-1 dark:bg-white/[.08]" > {/* Locked while a reply is arriving, which is still being read from the endpoint it was sent to. */} {(Object.keys(ENDPOINTS) as Endpoint[]).map((path) => ( <button key={path} role="tab" aria-selected={endpoint === path} disabled={busy} onClick={() => setEndpoint(path)} className={`rounded-md px-3 py-1.5 font-mono text-sm disabled:opacity-50 ${ endpoint === path ? "bg-white text-black shadow-sm dark:bg-zinc-800 dark:text-zinc-50" : "text-zinc-600 dark:text-zinc-400" }`} > {path} </button> ))} </div> <p className="text-sm text-zinc-600 dark:text-zinc-400"> {ENDPOINTS[endpoint].summary}.{" "} <button onClick={() => learnMore.current?.showModal()} className="font-medium text-blue-600 underline-offset-2 hover:underline dark:text-blue-400" > Learn more </button> </p> {/* The box fills whatever space the window leaves, and the thread scrolls inside it, so a long conversation never pushes the input off the screen. */} <div className="flex min-h-0 w-full max-w-[650px] flex-1 flex-col overflow-hidden rounded-2xl border border-black/[.08] bg-white dark:border-white/[.12] dark:bg-black"> <div className="flex items-center justify-between gap-2 border-b border-black/[.08] px-3 py-2 dark:border-white/[.12]"> <select aria-label="Model" value={model} disabled={busy} onChange={(event) => setModel(event.target.value as Model)} className="rounded-md bg-transparent py-1 text-sm text-black disabled:opacity-50 dark:text-zinc-50" > {(Object.keys(MODELS) as Model[]).map((id) => ( <option key={id} value={id}> {MODELS[id]} </option> ))} </select> <button onClick={confirmStartOver} disabled={locked || messages.length === 0} className="rounded-md px-2 py-1 text-sm text-blue-600 disabled:opacity-40 dark:text-blue-400" > Start new chat </button> </div> <div ref={thread} className="flex flex-1 flex-col gap-2 overflow-y-auto p-4"> {messages.map((turn, index) => { const mine = turn.role === "user"; const typing = busy && index === messages.length - 1; return ( <div key={index} className={`flex max-w-[80%] flex-col gap-1 ${mine ? "items-end self-end" : "items-start self-start"}`} > {/* Claude answers in Markdown, so its bubbles render it: `prose` styles the lists, headings, and code blocks it produces. A prompt is shown as typed, and `whitespace-pre-wrap` keeps its newlines, which HTML would otherwise collapse. A `<div>`, not a `<p>`, because a list or a code block cannot sit inside a paragraph. */} <div className={`min-w-0 rounded-2xl px-4 py-2 leading-6 ${ mine ? "whitespace-pre-wrap rounded-br-md bg-blue-500 text-white" : "prose prose-zinc max-w-none rounded-bl-md bg-zinc-100 text-black dark:prose-invert dark:bg-zinc-800 dark:text-zinc-50" }`} > {mine ? ( turn.content ) : ( <Markdown remarkPlugins={[remarkGfm]} components={{ a: NewTabLink }} > {turn.content} </Markdown> )} {typing && ( <span className="ml-0.5 inline-block h-4 w-2 animate-pulse bg-current align-middle" /> )} </div> {typing && activity && ( <span className="animate-pulse text-xs text-zinc-500 dark:text-zinc-400"> {activity}… </span> )} {turn.error && ( <span className="flex gap-2 text-xs"> <span role="alert" className="text-red-600 dark:text-red-400"> Not delivered </span> <button onClick={() => retry(index)} disabled={locked} className="font-medium text-blue-600 disabled:opacity-40 dark:text-blue-400" > Try again </button> </span> )} </div> ); })} </div> {/* The input is never disabled. A disabled input drops focus, so the cursor would leave it on every send; only the button locks. */} <form onSubmit={send} className="flex gap-2 border-t border-black/[.08] p-3 dark:border-white/[.12]" > <input ref={input} name="prompt" placeholder="Ask Claude something" autoComplete="off" autoFocus className="min-w-0 flex-1 rounded-full border border-black/[.08] px-4 py-2 text-black placeholder:text-zinc-400 dark:border-white/[.12] dark:text-zinc-50" /> <button disabled={locked} className="rounded-full bg-black px-4 py-2 font-medium text-white disabled:opacity-50 dark:bg-zinc-50 dark:text-black" > {busy ? "…" : "Send"} </button> </form> </div> <Modal ref={learnMore} title={endpoint} onClose={() => input.current?.focus()}> {ENDPOINTS[endpoint].explanation} </Modal> <Modal ref={errorModal} title="Not delivered" onClose={() => input.current?.focus()}> <p>{problem}</p> <p>Your message is still in the thread. Use Try again to send it again.</p> </Modal> {/* The button that closed the dialog leaves its `value` in `returnValue`. Cancel, Escape, and a backdrop click leave it empty. */} <Modal ref={startOverModal} title="Start new chat?" confirm="Start new chat" onClose={(event) => event.currentTarget.returnValue === "confirm" ? startOver() : input.current?.focus() } > <p> This conversation will be deleted from the server. You cannot get it back. </p> </Modal> </div> ); } // Links in a reply, like the sources under a web search, open in a new tab so // the chat stays where it is. `noreferrer` keeps the opened page from reaching // back into this one through `window.opener`. function NewTabLink({ href, children }: React.ComponentProps<"a">) { return ( <a href={href} target="_blank" rel="noreferrer"> {children} </a> ); } // `<dialog>` with `showModal()` does the hard parts of a modal on its own: it // sits above the page, blocks clicks behind it, keeps focus inside, and closes // on Escape. A `<form method="dialog">` button closes it with no handler. // Pass `confirm` to turn Close into Cancel plus a button that says yes. function Modal({ ref, title, confirm, onClose, children, }: { ref: React.Ref<HTMLDialogElement>; title: string; confirm?: string; onClose: (event: React.SyntheticEvent<HTMLDialogElement>) => void; children: React.ReactNode; }) { return ( <dialog ref={ref} onClose={onClose} onClick={(event) => { // A click on the backdrop if (event.target === event.currentTarget) { // lands on the dialog itself, event.currentTarget.close(); // not on anything inside it. } }} className="m-auto w-[calc(100%-2rem)] max-w-md rounded-2xl bg-white p-0 text-black backdrop:bg-black/40 dark:bg-zinc-900 dark:text-zinc-50" > <div className="p-6"> <h2 className="font-mono text-lg font-semibold">{title}</h2> <div className="mt-3 flex flex-col gap-3 text-sm leading-6 text-zinc-700 dark:text-zinc-300"> {children} </div> <form method="dialog" className="mt-5 flex justify-end gap-2"> {confirm ? ( <> <button className="rounded-full px-4 py-2 text-sm font-medium text-zinc-700 dark:text-zinc-300"> Cancel </button> <button value="confirm" className="rounded-full bg-red-600 px-4 py-2 text-sm font-medium text-white" > {confirm} </button> </> ) : ( <button className="rounded-full bg-black px-4 py-2 text-sm font-medium text-white dark:bg-zinc-50 dark:text-black"> Close </button> )} </form> </div> </dialog> ); }

Run the Drash App

npm run dev

Then open http://localhost:3000 . The chat is the application’s home page; app/page.tsx replaces the placeholder page that create-next-app generates.

Verification

With npm run dev running on localhost:3000:

Using the Browser

Open http://localhost:3000 , then:

  1. You should be on the /api/chat/stream tab. Send a prompt. The reply should grow a few words at a time.
  2. Switch to the /api/chat tab and send “shorter”. The reply should arrive all at once and build on the first one, because both endpoints share the conversation.
  3. Reload the page. The conversation should come back.
  4. Click “Start new chat” and confirm. The thread should clear and stay empty after a reload.
  5. Switch back to /api/chat/stream and ask “What is the latest version of Deno?” “Searching the web…” should appear while Claude searches, and the reply should end with a Sources list whose links open in a new tab.
  6. Paste a URL and ask for a summary. “Reading a page…” should appear instead.

Using cURL

-c jar -b jar makes curl save the cookie and send it back, the way a browser does:

curl -c jar -b jar -X POST localhost:3000/api/chat -d '{"prompt":"My name is Eric"}' -> {"text":"..."} (and set-cookie: conversation=...) curl -c jar -b jar -N -X POST localhost:3000/api/chat/stream -d '{"prompt":"What is my name?"}' -> data: {"text":"Eric"} ... event: done curl -b jar localhost:3000/api/conversation -> {"messages":[{"role":"user","content":"My name is Eric"}, ... ]} curl -b jar -X DELETE localhost:3000/api/conversation -> 204 curl -b jar localhost:3000/api/conversation -> {"messages":[]} curl -X POST localhost:3000/api/chat -d 'not json' -> 400 Body must be JSON curl -X POST localhost:3000/api/chat -d '{}' -> 400 Body must have a `prompt` string curl -X POST localhost:3000/api/chat -d '{"prompt":"hi","model":"gpt-9"}' -> 400 Unknown model curl -X POST localhost:3000/api/chat -d '{"prompt":"hi","model":"claude-haiku-4-5"}' -> {"text":"..."} curl -X GET localhost:3000/api/chat -> 501 Not Implemented curl -X POST localhost:3000/api/nope -> 404 Not Found

The second request shows the memory working: it sends one question and no history, yet the answer depends on the first request. Drop -b jar from it and the server starts a new conversation, with no idea who you are.

-N on the second request turns off curl’s buffering. Without it, you still get every frame, but they all arrive at once at the end, which makes a working stream look broken.

The last two show why the catch-all route is worth having: both responses come from the chain, for URLs Next.js passed through without complaint. The 501 appears because route.ts exports GET for /api/conversation — without that export, Next.js would answer 405 before the chain saw the request.

Proving the Streaming Caveat

Point the application at a key you know is wrong and send the same prompt to both endpoints. The difference between the two responses is why the warning in Handling Errors From the API exists:

curl -X POST localhost:3000/api/chat -d '{"prompt":"hi"}' -> 500 Internal Server Error The server's Anthropic API key was rejected. curl -N -X POST localhost:3000/api/chat/stream -d '{"prompt":"hi"}' -> 200 OK, content-type: text/event-stream event: error data: {"message":"The server's Anthropic API key was rejected."}

Same failure, two shapes. The non-streaming endpoint had not sent its status line yet, so anthropicErrorToHTTPError turned a bad key into a 500. The streaming endpoint had already sent 200 before it ever called Claude, so all it can do is report the error in the body — and a client that ignores the error event sees a successful, empty answer.

Send a prompt from the page with the same wrong key, on either tab, and the prompt is marked “Not delivered” while the error modal says the key was rejected. Put the right key back, press “Try again”, and the same prompt goes through. The streaming tab gets there because the component acts on the error event. Delete that branch and the page shows an empty reply bubble and no error, because as far as it can tell, the response succeeded.

The Route Handler

app/api/[...drash]/route.ts is a catch-all route: the brackets and the ... mean it answers everything beneath /api, whatever the path. That is deliberate. Give Drash one segment and Next.js has already done the routing; give it the whole /api space and the chain decides what matches, which is the part you came here for. (Doubling the brackets — [[...drash]] — makes the segment optional too, so bare /api also reaches the chain instead of Next.js answering it. Nothing here needs that.)

Two consequences are easy to get wrong.

paths holds the URL the browser asks for, not a path relative to the handler. The file sits at app/api/[...drash]/, so the resources claim /api/chat, /api/chat/stream, and /api/conversation. A resource with paths = ["/chat"] never matches, and every request gets a 404.

Next.js dispatches on the export name. A route handler is a module exporting one function per HTTP method, and a method with no export is answered by Next.js with a 405 before your code runs. So every method any resource implements needs an export: GET and DELETE for /api/conversation, POST for the two chat endpoints. The exports are per file, not per resource, and that has a useful side effect: because GET is exported, GET /api/chat reaches the chain and gets the 501 Not Implemented that the base Resource returns for a method you did not define, rather than a 405 from Next.js. The chain decides.

All three exports call the same handle(), so there is one application and one .catch() whichever method the request uses.

Route handlers are not cached. GET handlers were cached by default in Next.js 14; that changed in Next.js 15 , and they now run on every request. Nothing needs export const dynamic = "force-dynamic" to stream correctly.

Why the Polyfill Entry Point

Every other example picks its entry point from one runtime’s URLPattern support. This one cannot, because Next.js runs in many places: the same route.ts runs on Node 20 and 22, which have no global URLPattern; on Node 24 and later, which do; and under Deno, Bun, or Cloudflare Workers through an adapter. The polyfill is the one import that works on all of them.

There is no Edge variant to weigh against it. Next.js 16.3 dropped runtime = 'edge'; a route that still declares it runs on Node anyway.

Choosing the Model

The page offers a choice of models, and the server has to decide which ones it will accept. app/models.ts holds the list, and both sides import it:

export const MODELS = { "claude-opus-5": "Claude Opus 5", "claude-sonnet-5": "Claude Sonnet 5", "claude-haiku-4-5": "Claude Haiku 4.5", "claude-fable-5-1": "Claude Fable 5.1", } as const;

The component renders it as a <select> and sends the chosen ID with each prompt. The route handler checks it with isModel before anything else happens:

const model = body.model ?? DEFAULT_MODEL; if (!isModel(model)) { throw new HTTPError(Status.BadRequest, "Unknown model"); }

The check is not optional just because the page only offers valid IDs. The endpoint is a public URL, and a caller can send any string — including the ID of a model that costs more than the ones you chose to offer. One list shared by both sides means the menu and the check can never disagree, and it is safe to import into a client component because it is plain data with no secrets in it.

The model is picked per request, not per conversation. Switch models partway through and the next reply comes from the new one, reading the same stored history.

Remembering the Conversation

The Claude API keeps no state between calls. Every request has to carry the whole conversation, or the model answers a follow-up with no idea what came before it. In this example, the server holds that history:

const conversations = new Map<string, Turn[]>();

One list of turns per conversation, keyed by an ID. The browser holds only the ID, in a cookie, and sends it back with every request without being asked. openConversation reads it, or makes a new one when there is none:

function openConversation(request: Request) { const cookies = request.headers.get("cookie") ?? ""; const found = cookies.match(new RegExp(`(?:^|;\\s*)${COOKIE}=([^;]+)`))?.[1]; const id = found ?? crypto.randomUUID(); const headers: Record<string, string> = found ? {} : { "set-cookie": `${COOKIE}=${id}; Path=/; HttpOnly; SameSite=Lax` }; return { id, turns: conversations.get(id) ?? [], headers }; }

A resource is handed a Web Request, so the cookie is just a header, and setting one is just a header on the Response. The resources spread headers into their response, and only the first response of a conversation carries set-cookie. HttpOnly keeps the ID out of reach of page scripts; add Secure once the application is served over HTTPS.

Each chat resource sends Claude the stored turns plus the new prompt, then calls remember once the reply is complete:

function remember(id: string, turns: Turn[], prompt: string, reply: string) { if (reply === "") { return; } conversations.set(id, [ ...turns, { role: "user", content: prompt }, { role: "assistant", content: reply }, ]); }

The prompt and the reply are stored together or not at all. A request that fails leaves the history as it was, so the next request never sends Claude a question that was never answered. The empty-reply guard exists for the same reason: the API rejects a turn with no text, and storing one would make every later request in that conversation fail.

Conversation is the third resource, and it is what makes the memory visible. GET /api/conversation returns the stored turns, which the page loads when it opens; DELETE forgets them, which is what the page’s “Start new chat” button calls.

A Map is memory, not storage. Restarting the server forgets every conversation, and so does editing route.ts under npm run dev, which reloads the module. A serverless deployment is worse: each instance has its own memory, so consecutive requests can land on instances that have never heard of the conversation. Anything real keeps these lists in a database or a key-value store. The Map is the part to replace; the cookie, openConversation, and remember stay as they are. The history also grows with every turn, and so does each request’s cost, because the model reads all of it every time. A long-lived conversation eventually wants a cap, whether that means dropping the oldest turns or summarizing them.

The Non-Streaming Endpoint

Chat is the shorter of the two chat resources, and nothing in it goes beyond an ordinary resource method.

class Chat extends Resource { public paths = ["/api/chat"]; public async POST(request: Request) { const { prompt, model } = await readRequest(request); const { id, turns, headers } = openConversation(request); const message = await client.messages.create({ model, max_tokens: 16000, tools: webTools(model), messages, }); // ... remember(id, turns, prompt, text); return Response.json({ text }, { headers }); } }

The call sits inside a loop for a reason covered in Browsing the Web; when Claude does not browse, the loop runs once. Two details are worth pointing out.

content is a list of blocks, not a string. A reply can contain reasoning blocks, web searches and their results, and text, so message.content is a mixed array. Check each block’s type before reading .text:

for (const block of message.content) { if (block.type === "text") { text += resumes(previous, text) + block.text; } previous = block.type; }

In TypeScript that check is also what narrows the type — block.text does not compile on the unchecked union.

max_tokens is a ceiling, not a target. 16,000 is a reasonable default for a request that waits for the full reply: large enough that answers are not cut off mid-sentence, small enough that the request finishes inside the SDK’s ten-minute timeout. A reply that hits the ceiling comes back with stop_reason set to "max_tokens" and no indication in the text itself, so check stop_reason if truncation matters to you.

The Streaming Endpoint

ChatStream returns before the model has finished. The body is a ReadableStream that fills in as events arrive. Stripped of browsing, which has its own section, it comes down to this:

const body = new ReadableStream({ async start(controller) { let reply = ""; const stream = client.messages.stream({ model, max_tokens: 64000, tools: webTools(model), messages, }); for await (const event of stream) { if ( event.type === "content_block_delta" && event.delta.type === "text_delta" ) { reply += event.delta.text; controller.enqueue(encoder.encode( `data: ${JSON.stringify({ text: event.delta.text })}\n\n`, )); } } remember(id, turns, prompt, reply); controller.close(); }, }); return new Response(body, { headers: { ...headers, "content-type": "text/event-stream" }, });

The text is collected into reply as it goes out, because the server has to remember the answer it just streamed and nothing else will hand it back. remember runs after the loop, so a stream that fails partway never stores half an answer. The set-cookie header rides along with the others: headers go out before the first frame, which is exactly when a new conversation’s cookie needs to reach the browser.

max_tokens goes up to 64,000 here because the timeout that limited it no longer applies: a streaming connection does not wait for a complete response, so the model gets more room.

The stream carries several event types; content_block_delta with a text_delta is the one that holds generated text. The full sample also reads block starts and citations, which browsing needs; everything else — block stops, message-level usage totals — is skipped.

The x-accel-buffering: no header in the full sample is for deployment rather than for Next.js. A reverse proxy in front of your application may hold a response until it looks complete, which turns a working stream into one long pause; that header tells nginx and the proxies that copied its convention not to.

The Frame Format

Each data: line is one server-sent event, and the blank line is what ends it. A browser reads these with EventSource, or with fetch if you need to POST:

data: {"text":"Let me look that up."} event: activity data: {"text":"Searching the web"} data: {"text":"\n\nThe latest release is"} event: done data: {}

An unnamed event is text to add to the reply. The named done event exists because otherwise a closed connection and a finished reply look identical to the client, and activity says what Claude is doing while it is not writing. The page treats any event it does not recognize as reply text, so every named event needs a matching branch in the client.

Expect a pause before the first token. Claude Opus 5 thinks before it answers, and the thinking is not sent by default — the stream carries thinking blocks with empty text, then starts producing text_delta events. If you want to show the reader something during that gap, ask for a summary of the reasoning and forward it on its own event:

client.messages.stream({ model: "claude-opus-5", max_tokens: 64000, thinking: { type: "adaptive", display: "summarized" }, messages: [...turns, { role: "user", content: prompt }], });

The deltas then arrive as thinking_delta on the same content_block_delta event, alongside the text_delta ones.

Browsing the Web

Claude can search the web and read pages, but only with tools it is given. Both endpoints pass the same two:

function webTools(model: Model): Anthropic.Messages.ToolUnion[] { if (model === "claude-haiku-4-5") { return [ { type: "web_search_20250305", name: "web_search", max_uses: 5 }, { type: "web_fetch_20250910", name: "web_fetch", max_uses: 5 }, ]; } return [ { type: "web_search_20260209", name: "web_search", max_uses: 5 }, { type: "web_fetch_20260209", name: "web_fetch", max_uses: 5 }, ]; }

These are server tools: Anthropic runs them, not you. Claude decides on its own when a prompt needs the web, the search or fetch happens on Anthropic’s side, and the reply comes back with the results already read. Nothing in the route handler fetches a URL, parses HTML, or calls a search API, and there is no tool-call loop to write. web_search finds pages; web_fetch reads one page in full, but only from a URL already in the conversation (in your prompt or an earlier search result), so it cannot be steered to an address the model made up.

The versions differ by model. The _20260209 versions let Claude filter results with code before reading them, which keeps irrelevant pages out of the context and spends fewer tokens; Claude Haiku 4.5 does not support them and uses the older ones instead. max_uses caps searches and fetches per reply. Each search is billed on top of tokens, and each result is read as input, so without a cap one curious prompt can run up the cost.

Three things change once a reply can browse.

A reply can pause. Browsing runs as a loop on Anthropic’s side, and a long one stops partway with stop_reason: "pause_turn" rather than finishing. Sending the reply so far back as an assistant turn picks it up where it stopped:

if (message.stop_reason !== "pause_turn" || round === MAX_CONTINUATIONS) { break; } messages.push({ role: "assistant", content: message.content });

Push message.content whole, not just its text: the search blocks at the end of it are how the API knows what to resume. Do not add a user message saying “continue” either; the trailing search block already says that. MAX_CONTINUATIONS stops a runaway loop from billing forever.

Text arrives in pieces. Claude often writes a sentence, searches, then keeps writing, so the text is split across blocks with search results in between. Joined as-is, “Let me check.” runs straight into the answer. resumes puts a paragraph break wherever text starts after a non-text block. Adjacent text blocks are left alone, because those are one sentence split up by citations.

Results must name their sources. Text that draws on a search result carries a citation with the page’s URL and title. Anthropic’s terms require showing those wherever web search results reach users, so both endpoints collect them (from block.citations, or from citations_delta events while streaming) and listSources adds them under the reply as a Markdown list. The page renders it like any other Markdown, and NewTabLink opens each link in a new tab so the chat stays put.

Only the reply’s text is remembered, sources included. The search results themselves are not, so a follow-up question that needs them has Claude search again. Storing message.content whole would avoid that, at the cost of sending every page it read on every later request.

The streaming endpoint has one more job: a search takes several seconds with no text to show. When a server_tool_use block starts, it sends an activity event saying “Searching the web” or “Reading a page”, and the page shows that under the bubble until text resumes.

The React Client

The browser half is one component marked "use client", which is what makes it run in the browser rather than on the server. It is the only file here that ships JavaScript to the browser, and it never touches the API key — it calls your endpoint, and your endpoint calls Anthropic.

One Component, Two Endpoints

The two tabs at the top of the page pick which resource the next prompt goes to. The tab only changes the URL: both resources take the same { prompt } body and add to the same stored conversation, so one conversation can move between them, and the reply lands in the same bubble either way. The tabs are disabled while a reply is arriving, because that reply is still being read from the endpoint it was sent to.

Each tab’s caption has a “Learn more” button that opens a modal explaining that endpoint: what the server sends, what the client does with it, and what happens when something fails. The explanations live next to the endpoint paths, in the same ENDPOINTS object the tabs are built from, so adding a tab means writing its explanation in the same place.

/api/chat needs almost nothing from the client. It answers once, with JSON, so the reply is one response.json() and one append:

if (endpoint === "/api/chat") { const { text } = await response.json(); append(text); setBusy(false); return; }

Try it on a prompt that takes a while to answer. The pulsing block sits alone in an empty bubble for the whole generation, then the entire reply appears at once. That wait is the reason /api/chat/stream exists, and the rest of this section covers what it costs the client to avoid it.

Reading the Stream

The obvious tool for server-sent events is EventSource, and it is the wrong one here: EventSource issues a GET and cannot send a body, while the prompt has to be POSTed. So the response body is read directly:

const response = await fetch(endpoint, { method: "POST", headers: { "content-type": "application/json" }, body: JSON.stringify({ prompt }), }); const reader = response.body.getReader(); const decoder = new TextDecoder();

Reassembling the Frames

reader.read() resolves with whatever bytes have arrived, which has nothing to do with where frames begin and end. One read can deliver three frames, or half of one. Treating each read as a frame produces a parser that works until the first time the network splits a message — which is to say, it works in development and fails in production.

The fix is three lines:

buffer += decoder.decode(value, { stream: true }); const frames = buffer.split("\n\n"); buffer = frames.pop() ?? "";

Everything before a blank line is a complete frame. Whatever follows the last blank line is an unfinished one, so pop() takes it back out and leaves it in the buffer for the next read to complete. { stream: true } does the same job one level down, holding back a multi-byte character that got split across reads instead of decoding it into a replacement character.

Each complete frame is then read for its event name and its payload:

const name = frame.match(/^event: (.*)$/m)?.[1] ?? "message"; const data = frame.match(/^data: (.*)$/m)?.[1];

A frame with no event: line is a plain message, which is where the text arrives. The other names are the ones the endpoint sends deliberately: activity says what Claude is doing while it is not writing, done means the reply finished, and error means it did not. Handling error is not optional — see below for what ignoring it looks like.

Showing the Conversation

The server holds the conversation; the component holds a copy of it to draw. That copy is one piece of state, a list of turns in the shape the server stores them, rendered like a messaging app: your prompts on the right, Claude’s replies on the left, oldest at the top.

type Turn = { role: "user" | "assistant"; content: string; error?: string; };

When the page opens, the copy is filled from the server, so a reload shows the conversation where you left it:

useEffect(() => { fetch("/api/conversation") .then((response) => response.json()) .then((body: { messages: Turn[] }) => setMessages(body.messages)) .catch(() => {}) .finally(() => setLoaded(true)); }, []);

Sending stays locked until loaded is set. Otherwise a prompt sent before the stored thread arrived would be overwritten by it.

The cookie never appears in this code. The browser attaches it to every same-origin fetch on its own, and HttpOnly means the component could not read it if it tried. As far as the component knows, it sends a prompt and the server somehow remembers.

Sending a prompt clears the input, as a messaging app does, and appends two turns to the screen: the prompt, and an empty assistant turn for the reply to land in. The request itself carries only the prompt.

setMessages((turns) => [ ...turns, { role: "user", content: prompt }, { role: "assistant", content: "" }, ]);

Both endpoints write the reply into that last turn through one helper, append. The whole-reply tab calls it once; the streaming tab calls it once per frame. It uses the function form of setMessages:

const append = (text: string) => { setMessages((turns) => { const last = turns[turns.length - 1]; return [ ...turns.slice(0, -1), { ...last, content: last.content + text }, ]; }); };

Reading the previous list from the callback rather than from messages keeps the append correct when several frames land between renders, which they will. messages in that closure still holds the list as it was when the prompt was sent, so using it would make each frame overwrite the one before.

When a Message Is Not Delivered

A failed request removes the half-written reply and stores the reason on the prompt’s error, which renders as “Not delivered” under its bubble. That matches the server exactly: remember never ran, so the prompt exists only on screen, and the next request carries on from the last turn that succeeded. Reload the page and the failed prompt is gone, because the server never had it.

The same failure also opens a modal with the reason. Every failure has a reason, and the reasons come from three places:

Where it failedWhere the reason comes from
The request never reached the serverThe component: fetch rejected, so it writes its own
The server answered with an error statusThe response body, which is the HTTPError message
The stream reported an error after 200The message in the event: error frame

The server writes the last two, which is why anthropicErrorToHTTPError gives every HTTPError a message a person can act on (see Handling Errors From the API). A stream that ends without a done event counts as a failure too, so a dropped connection is not mistaken for a short reply.

“Try again” sends the same prompt again. Because the server never stored the prompt, retrying takes only this:

setMessages((turns) => turns.filter((_, i) => i !== index)); submit(prompt);

The failed turn is removed, and submit appends the prompt at the end, like any new prompt. Both updates use the function form, so they apply in order even though they are queued in the same tick.

“Start new chat” asks first, because the server forgets the conversation for good. Only after you confirm does the page call DELETE /api/conversation and then clear the screen. The cookie stays, so the next prompt starts a fresh history under the same ID.

What the Styling Is Doing

The classes are Tailwind’s, which create-next-app --tailwind set up — it installs Tailwind CSS 4, adds @import "tailwindcss" to app/globals.css, and generates a theme block. There is no tailwind.config.js to edit; version 4 configures itself from CSS. The palette deliberately matches the home page create-next-app generates, so the two pages look like one application.

The bubbles are where the styling stops being decoration, because the two sides hold different kinds of text:

{mine ? ( turn.content ) : ( <Markdown remarkPlugins={[remarkGfm]}>{turn.content}</Markdown> )}

Claude answers in Markdown. Printed as-is, a reply is full of **, #, and backticks, and a numbered list collapses into one long paragraph, because HTML folds runs of whitespace into a single space. That looks like a bug in your streaming code, but it is not. <Markdown> turns the text into React elements (paragraphs, lists, headings, code blocks), and remark-gfm adds the GitHub extensions models use constantly: tables, strikethrough, and task lists. prose from the typography plugin gives those elements spacing and type, prose-zinc matches the rest of the palette, and dark:prose-invert flips it in dark mode. max-w-none lifts the reading-width cap prose sets, so the bubble’s own max-w-[80%] is the only limit.

react-markdown builds elements rather than an HTML string, so there is no dangerouslySetInnerHTML and no sanitizer to configure: raw HTML in a reply is ignored rather than rendered, and a model that echoes a <script> tag back cannot run it. During streaming, <Markdown> re-renders on every append, so an unfinished code fence shows as a code block that fills in as the rest arrives.

Your own prompts are not rendered as Markdown. They are shown as typed, with whitespace-pre-wrap keeping the newlines you entered and still wrapping long lines. Because a list or a code block cannot sit inside a paragraph, the bubble is a <div> rather than a <p>; React would otherwise warn about invalid nesting.

One smaller thing is worth knowing, because it is a property of the generated project rather than a choice this page made.

font-sans is not redundant. The generated layout.tsx loads Geist and exposes it as --font-geist-sans, but the generated globals.css then sets body { font-family: Arial, Helvetica, sans-serif }, which wins. font-sans is what reaches past that to the font the layout actually loaded. The generated home page does the same thing, for the same reason.

The chat box grows and shrinks with the window. The page’s outer element is h-dvh, exactly as tall as the visible window, and every element between it and the box is a flex column with flex-1, so each one takes the height its parent has left after the heading, tabs, and summary line. min-h-0 on each of them is what makes that work: a flex item will not shrink below the height of its content by default, so without it a long conversation would stretch the box past the window instead of scrolling inside it. dvh rather than vh matters on phones, where the browser’s address bar slides in and out and changes how much of the page is visible. The width follows the window the same way: each of those elements is w-full, because items-center on a parent otherwise shrinks a child to the width of its content. The box stops growing at max-w-[650px], because lines much longer than that get hard to read, and the parent’s items-center keeps it centered once it does.

The box’s own layout is what keeps the input in reach. The box is a flex column: the thread takes the leftover height with flex-1 and scrolls with overflow-y-auto, and the form sits under it at its natural height. A long conversation scrolls inside the box rather than pushing the input off the screen. For the same reason, the thread follows new bubbles by scrolling itself with scrollTo rather than calling scrollIntoView, which would scroll the page too.

Everything else follows one boolean:

const [busy, setBusy] = useState(false);

It disables the Send button, the tabs, the model selector, “Try again”, and “Start new chat”, so nothing else can be sent while a reply is open. It also changes the button’s label and renders a pulsing block at the end of the newest bubble. Before the first token arrives, that bubble is empty, so the block stands in for a typing indicator. It is the only sign that more text is coming, which matters most during the pause before the first token — see the callout above for why that pause exists — and on the /api/chat tab, where the pause lasts the whole reply.

Handling Errors From the API

Error handling covers how a throw reaches your .catch(). What is specific here is the translation step in front of it: a failure talking to Anthropic is a failure of your server, so the status you return is rarely the status you received:

The SDK throwsYou throwWhy
Anthropic.AuthenticationErrorStatus.InternalServerErrorYour key is wrong or missing. The caller did nothing to cause it and can do nothing about it.
Anthropic.RateLimitErrorStatus.TooManyRequestsThe one case where passing the status through is right — your caller should back off too.
Anthropic.NotFoundErrorStatus.BadRequestThe chosen model is not available to your key. The caller chose it, so it is theirs to change.
Anthropic.BadRequestErrorStatus.BadRequestAlmost always the prompt, which came from the caller.
Anthropic.APIErrorStatus.BadGatewayThe parent class, so it catches outages and timeouts. It must be checked last.

anthropicErrorToHTTPError in the sample does exactly that and nothing else. Returning the unrecognized error unchanged matters: a TypeError in your own code is a bug, and turning it into a tidy 502 hides it.

Each HTTPError also carries a message, such as “The server’s Anthropic API key was rejected.” handle() sends that message as the response body, and the page shows it in its error modal. Without one, the body would be empty and the page would have nothing to tell you but a number.

A refusal is the one failure that is not an exception. When Claude declines a prompt the API answers 200 with stop_reason: "refusal", so both resources check for it and throw an HTTPError of their own: Chat after the reply, ChatStream when a message_delta event carries that stop reason.

None of this applies once the stream is open. The status line and headers go out with the first byte. After that a failure cannot become a 500, because a 200 has already been sent — the client has a successful response containing half an answer.

Two habits follow. Validate the request body and construct the client before you return the Response, so anything the caller got wrong still produces a real status code. And inside the stream, report failures in-band:

const failure = anthropicErrorToHTTPError(error); const message = failure instanceof HTTPError ? failure.message : "The reply stopped partway through."; send(`event: error\ndata: ${JSON.stringify({ message })}\n\n`);

The same translation runs here, so a bad key reads the same on both tabs; only the delivery differs. Clients have to handle that event for it to mean anything, which is why the component above returns early on it rather than waiting for a done that is never coming.

Getting the API Key to the Resource

The key is read once, when the module first loads, from process.env.ANTHROPIC_API_KEY, and one client is built that both resources close over. Next.js loads .env.local from the project root on its own, so a key kept there needs no flag and no loader.

The name matters. Next.js exposes an environment variable to the browser only if its name begins with NEXT_PUBLIC_, and inlines it into the client bundle when it does. ANTHROPIC_API_KEY has no such prefix, so it stays on the server — and renaming it to NEXT_PUBLIC_ANTHROPIC_API_KEY would publish your key to every visitor.

Where to Next

  • Calling Claude on Node — the two endpoints without a framework, where the resource is handed a context object you build instead of a Request
  • Deploy a Basic App to Vercel (api/) — the same polyfill entry point in a Vercel Function, without Next.js in front of it
  • Creating Middleware — put an API key check or a request log in front of both endpoints
  • Rate Limiter — the bundled middleware for capping how often a caller can reach an endpoint this expensive
  • Grouping Resources — mount both under /api/v1 and share middleware between them

The conversation here lives in the server’s memory, which is enough to survive a reload and not enough to survive a restart. Moving it into a database means replacing only the Map, and the choice of database is yours, not Drash’s.

The finished app is in the repository at examples/frameworks/next-js/calling-claude.

Last updated on