Skip to Content
ExamplesRuntimesDenoCalling Claude

Calling Claude on Deno

Two endpoints backed by the Claude API — one that answers all at once, one that streams

Overview

A resource calls Claude the way it does anything else: an ordinary await inside POST.

What does need thought is the shape of the answer. A model generates its reply a piece at a time, and a request that waits for the whole thing can sit for a minute before anything reaches the browser. So this example builds both:

EndpointAnswers with
POST /chatOne JSON body, once the model has finished
POST /chat/streamA server-sent event  per piece, as the model produces it

Your resource returns whatever your runtime wants back, and Drash passes that value through untouched. That is what makes a streamed body a non-event for the framework: the resource below returns a Response, and a streamed one is a Response whose body is a ReadableStream.

Objectives

To gain familiarity with:

  • reading and validating a JSON request body inside a resource;
  • returning a Response whose body is a ReadableStream;
  • turning an API client’s errors into ones the chain can render; and
  • reading a secret on a runtime that asks permission for it.

Steps

Folder Structure End State

    • app.ts

Set Up the Project

You need the runtime itself — the quickstart sets one up if you do not have one — and an Anthropic API key .

# No install step. Deno resolves both imports on first run. export ANTHROPIC_API_KEY="sk-ant-..."

Keep the key on the server. Every line below runs in your application, never in a browser. An Anthropic key in client-side code is readable by anyone who opens devtools, and every request they make with it is charged to your account.

Create the Following File

app.ts
import Anthropic from "npm:@anthropic-ai/sdk"; import { Application, HTTPError, Resource, } from "jsr:@drashland/drash/modules/http.native"; import { Status } from "jsr:@drashland/drash/core/http/response/Status"; const MODEL = "claude-opus-5"; const client = new Anthropic({ // `--allow-env` is what lets apiKey: Deno.env.get("ANTHROPIC_API_KEY"), // this read the key. }); // Both resources start here. A `Request` parses its own body, and this throws // rather than returning an error, which is how a resource reports anything the // caller got wrong. async function readPrompt(request: Request): Promise<string> { let body: { prompt?: unknown }; try { body = await request.json(); } catch { throw new HTTPError(Status.BadRequest, "Body must be JSON"); } if (typeof body.prompt !== "string" || body.prompt === "") { throw new HTTPError(Status.BadRequest, "Body must have a `prompt` string"); } return body.prompt; } // Anthropic's failures are not your caller's failures. Restate each one as the // status your caller should actually see. Order matters — the specific classes // come before `APIError`, which is their shared parent. function anthropicErrorToHTTPError(error: unknown): unknown { if (error instanceof Anthropic.AuthenticationError) { return new HTTPError(Status.InternalServerError); // Your key is wrong. } if (error instanceof Anthropic.RateLimitError) { return new HTTPError(Status.TooManyRequests); // Pass the limit on. } if (error instanceof Anthropic.BadRequestError) { return new HTTPError(Status.BadRequest); // Usually the prompt. } if (error instanceof Anthropic.APIError) { return new HTTPError(Status.BadGateway); // Upstream is down. } return error; // Not ours. Leave it. } class Chat extends Resource { // Answers with the whole reply. public override paths = ["/chat"]; public override async POST(request: Request) { const prompt = await readPrompt(request); let message; try { message = await client.messages.create({ model: MODEL, max_tokens: 16000, messages: [{ role: "user", content: prompt }], }); } catch (error) { throw anthropicErrorToHTTPError(error); } const text = message.content // `content` holds blocks of .filter((block) => block.type === "text") // several types. Keep the text .map((block) => block.text) // ones before reading `.text`. .join(""); return Response.json({ text }); // Whatever you return is what } // `app.handle()` resolves with. } class ChatStream extends Resource { // Answers a piece at a time. public override paths = ["/chat/stream"]; public override async POST(request: Request) { const prompt = await readPrompt(request); // Throws before the `Response` // exists — that matters. const stream = client.messages.stream({ model: MODEL, max_tokens: 64000, messages: [{ role: "user", content: prompt }], }); const encoder = new TextEncoder(); const body = new ReadableStream({ async start(controller) { const send = (frame: string) => controller.enqueue(encoder.encode(frame)); try { for await (const event of stream) { if ( event.type === "content_block_delta" && event.delta.type === "text_delta" ) { send(`data: ${JSON.stringify({ text: event.delta.text })}\n\n`); } } send("event: done\ndata: {}\n\n"); } catch (error) { // Too late for a status code. send( // Say so in the stream instead. `event: error\ndata: ${ JSON.stringify({ message: "upstream failed" }) }\n\n`, ); console.error(error); } finally { controller.close(); } }, }); return new Response(body, { // A streamed body is still just headers: { // a `Response`. Drash passes it "content-type": "text/event-stream", // through untouched. "cache-control": "no-store", }, }); } } const app = Application .builder() .resources(Chat, ChatStream) .build(); const hostname = "localhost"; const port = 1447; Deno.serve({ hostname, port, onListen: ({ hostname, port }) => { console.log(`\nDrash running at http://${hostname}:${port}`); }, handler: (request: Request): Promise<Response> => { return app .handle<Response>(request) .catch((error) => { // `name` is checked before `instanceof` because `instanceof` fails when // two copies of Drash end up in one bundle. if (error.name === "HTTPError") { return new Response(error.message, { status: error.status_code, statusText: error.status_code_description, }); } console.error(error); return new Response("The server could not generate a response", { status: 500, }); }); }, });

Run the Drash App

deno run --allow-net --allow-env app.ts

Deno asks for permission per capability. --allow-net is what lets the process reach api.anthropic.com and listen on a port; --allow-env is what lets it read ANTHROPIC_API_KEY. If the npm compatibility layer wants anything else, Deno stops and names the flag.

Or keep the key in a file. Deno never reads a .env on its own, so a key sitting in one reaches nothing and the client fails on its first request. --env-file is what loads it; --allow-env is still what permits reading it, so local development wants both: deno run --allow-net --allow-env --env-file app.ts.

Verification

With the application running on localhost:1447:

curl -X POST localhost:1447/chat -d '{"prompt":"Say hi"}' -> {"text":"..."} curl -N -X POST localhost:1447/chat/stream -d '{"prompt":"Hi"}' -> data: {"text":"..."} ... event: done curl -X POST localhost:1447/chat -d 'not json' -> 400 Body must be JSON curl -X POST localhost:1447/chat -d '{}' -> 400 Body must have a `prompt` string curl -X GET localhost:1447/chat -> 501 Not Implemented curl -X POST localhost:1447/nope -> 404 Not Found

-N on the second one turns off curl’s buffering. Without it you still get every frame, but they arrive in a lump at the end, which makes a working stream look like a broken one.

The last two come from the chain rather than from your code, on URLs your server was happy to hand over.

Proving the Streaming Caveat

Point the application at a key you know is wrong and send the same prompt to both endpoints. The difference is the whole reason the warning in Handling Errors From the API exists:

curl -X POST localhost:1447/chat -d '{"prompt":"hi"}' -> 500 Internal Server Error curl -N -X POST localhost:1447/chat/stream -d '{"prompt":"hi"}' -> 200 OK, content-type: text/event-stream event: error data: {"message":"upstream failed"}

Same failure, two shapes. The non-streaming endpoint still had its status line to spend, so anthropicErrorToHTTPError turned a bad key into a 500. The streaming endpoint had already sent 200 before it ever called Claude, so all it can do is say so in the body — and a client that ignores the error event sees a successful, empty answer.

The Non-Streaming Endpoint

Chat is the shorter of the two, and everything in it is an ordinary resource method.

class Chat extends Resource { public override paths = ["/chat"]; public override async POST(request: Request) { const prompt = await readPrompt(request); const message = await client.messages.create({ model: "claude-opus-5", max_tokens: 16000, messages: [{ role: "user", content: prompt }], }); // ... } }

Two things are worth naming.

content is a list of blocks, not a string. A reply can contain reasoning blocks, tool calls, and text, so message.content is a mixed array. Filter to the text blocks before reading .text:

const text = message.content .filter((block) => block.type === "text") .map((block) => block.text) .join("");

In TypeScript that filter is also what narrows the type — block.text does not compile on the unfiltered union.

max_tokens is a ceiling, not a target. 16,000 is a reasonable default for a request that waits for the full reply: large enough that answers are not cut off mid-sentence, small enough that the request finishes inside the SDK’s ten-minute timeout. A reply that hits the ceiling comes back with stop_reason set to "max_tokens" and no indication in the text itself, so check it if truncation would matter to you.

The Streaming Endpoint

ChatStream returns before the model has finished. The body is a ReadableStream that fills in as events arrive:

const stream = client.messages.stream({ model: "claude-opus-5", max_tokens: 64000, messages: [{ role: "user", content: prompt }], }); const body = new ReadableStream({ async start(controller) { for await (const event of stream) { if ( event.type === "content_block_delta" && event.delta.type === "text_delta" ) { controller.enqueue(encoder.encode( `data: ${JSON.stringify({ text: event.delta.text })}\n\n`, )); } } controller.close(); }, }); return new Response(body, { headers: { "content-type": "text/event-stream" }, });

max_tokens goes up to 64,000 here because the constraint it was working around is gone. A streaming connection does not time out waiting for a complete response, so the model gets room.

The stream carries several event types; content_block_delta with a text_delta is the one that holds generated text. Everything else — block starts and stops, message-level usage totals — is skipped by the if.

The Frame Format

Each data: line is one server-sent event, and the blank line is what ends it. A browser reads these with EventSource, or with fetch if you need to POST:

data: {"text":"Hello"} data: {"text":" there"} event: done data: {}

The named done event exists because a closed connection and a finished reply look identical to the client otherwise.

Expect a pause before the first token. Claude Opus 5 thinks before it answers, and the thinking is not sent by default — the stream carries thinking blocks with empty text, then starts producing text_delta events. If you want to show the reader something during that gap, ask for a summary of the reasoning and forward it on its own event:

client.messages.stream({ model: "claude-opus-5", max_tokens: 64000, thinking: { type: "adaptive", display: "summarized" }, messages: [{ role: "user", content: prompt }], });

The deltas then arrive as thinking_delta on the same content_block_delta event, alongside the text_delta ones.

What Your Runtime Expects Back

Deno.serve hands you a Web Request and expects a Response back, which is exactly what the chain takes and returns. A streamed body is just a Response whose body is a ReadableStream, so there is nothing to adapt: what the resource returns is what the server sends.

Handling Errors From the API

Error handling covers how a throw reaches your .catch(). What is specific here is the translation step in front of it: a failure talking to Anthropic is a failure of your server, so the status you return is rarely the status you received:

The SDK throwsYou throwWhy
Anthropic.AuthenticationErrorStatus.InternalServerErrorYour key is wrong or missing. The caller did nothing to cause it and can do nothing about it.
Anthropic.RateLimitErrorStatus.TooManyRequestsThe one case where passing the status through is right — your caller should back off too.
Anthropic.BadRequestErrorStatus.BadRequestAlmost always the prompt, which came from the caller.
Anthropic.APIErrorStatus.BadGatewayThe parent class, so it catches outages and timeouts. It must be checked last.

anthropicErrorToHTTPError in the sample does exactly that and nothing else. Returning the unrecognized error unchanged matters: a TypeError in your own code is a bug, and turning it into a tidy 502 hides it.

None of this applies once the stream is open. The status line and headers go out with the first byte. After that a failure cannot become a 500, because a 200 has already been sent — the client has a successful response containing half an answer.

Two habits follow. Validate the request body and construct the client before you return the Response, so anything the caller got wrong still produces a real status code. And inside the stream, report failures in-band:

send(`event: error\ndata: ${JSON.stringify({ message: "upstream failed" })}\n\n`);

Clients have to handle that event for it to mean anything, so document it alongside the endpoint.

Getting the API Key to the Resource

The key is read once, at startup, with Deno.env.get("ANTHROPIC_API_KEY"), and one client is built that both resources close over. Deno refuses that read without --allow-env, and it never loads a .env file unless --env-file tells it to.

Where to Next

  • Calling OpenAI — the same two endpoints against the Chat Completions API, where the streaming client fails earlier and more usefully
  • Calling Perplexity — the same again, answering with the sources the reply was built from
  • Creating Middleware — put an API key check or a request log in front of both endpoints
  • Rate Limiter — the bundled middleware for capping how often a caller can reach an endpoint this expensive
  • Grouping Resources — mount both under /api/v1 and share middleware between them

This example sends one prompt and gets one answer. Holding a conversation means sending the whole history on every request, since the API keeps no state between calls — the messages array grows, and where you store it is your decision, not Drash’s.

The finished app is in the repository at examples/runtimes/deno/calling-claude.

Last updated on