swarza

Host AI agents at a price that never moves.

An agent built with the AI SDK deploys like any other application, with one command. Time spent waiting on a model is part of your plan, and an agent that goes past its plan is throttled, never billed extra.

What an agent gets on Swarza now.

  • Streamed answers of any length
  • Work after the response, up to 15 minutes
  • Resumable chats, without Redis
  • Turso databases and S3-compatible buckets
  • MCP servers over Streamable HTTP
  • The AI SDK; Mastra and LangGraph.js as Node.js apps
  • Not yet: Durable workflows, for longer work (planned next)
  • Not yet: WebSockets (the AI SDK streams over HTTP)
  • Not yet: Python (Node.js 26 or Bun only)

The example agent’s code is in swarza/examples.

Work goes on after the response.

waitUntil and Next.js after() keep running for up to 15 minutes after the last request: saving a summary, calling a webhook, finishing a tool call. If work reaches the limit, the application’s logs say so.

// app/api/report/route.ts
import { after } from "next/server";

export async function POST(req: Request) {
  const chat = await req.json();
  after(() => summarize(chat));
  return Response.json({ queued: true });
}

A dropped connection picks up where it left off.

Follow the AI SDK’s resume guide: the page reconnects with useChat({ resume: true }) and gets the rest of the answer. The chunks wait in the application’s memory, so there is no Redis to run.

// components/chat.tsx
useChat({ id, resume: true });

// app/api/chat/[id]/stream/route.ts
const stream = await context.resumeExistingStream(id);
if (!stream) return new Response(null, { status: 204 });
return new Response(stream, { headers });

Conversations in a Turso database.

Create a database with one command and save each chat when its answer ends, so a reload, or a new deployment, shows the whole conversation.

$ swarza databases create notes --app my-app

import { connect } from "@tursodatabase/serverless";

const db = connect({
  url: process.env.DATABASE_URL,
  authToken: process.env.DATABASE_AUTH_TOKEN,
});
const notes = await db.all("SELECT * FROM notes");

Tools, MCP servers and agents on a schedule.

An MCP server over Streamable HTTP is an ordinary application. A scheduled job can run an agent every morning, for up to 5 or 15 minutes depending on the plan.

// swarza.json
{ "scheduledJobs": [
  { "schedule": "0 7 * * 1-5",
    "function": "jobs/digest.mjs" } ] }

// jobs/digest.mjs
export default async function () {
  const { text } = await generateText({ model, prompt });
  await sendDigest(text);
}

Answers stream for as long as the model needs.

While a request is open, its answer has no time limit. The AI SDK’s keepAliveMs sends a comment line every 15 seconds, so a model that thinks for minutes isn’t cut off on the way.
// app/api/chat/route.ts
export async function POST(req: Request) {
  const { messages } = await req.json();
  const result = streamText({ model, messages, tools });

  return createUIMessageStreamResponse({
    keepAliveMs: 15_000,
    stream: toUIMessageStream({ stream: result.stream }),
  });
}
0 MBof memory held by an idle application. While it sleeps, it isn’t billed.
180 msto wake an idle Next.js app on the server and answer, measured on staging.
15 minof work after the last request, for waitUntil and after().

Agent platforms side by side.

Where each one runs an agent, how you pay for it, and what happens to a long answer.

Add
Swarza compared with agent platforms
SwarzaVercelFunctionsCloudflareAgents
Entry price€5 / month (Starter)$20 / month per developer seat (Pro)$5 / month minimum (Workers Paid)
How you payOne flat monthly priceActive CPU, memory time and invocationsRequests and CPU time. Agents also by wall-clock time
Past the included usageThrottled, never billed extraBilled, uncapped. Pausing at a budget is opt-inBilled. Budget alerts don’t pause or cap usage
Longest streamed answerNo limit, with a keep-alive every 15 s800 s (1,800 s in beta)No limit while connected
Work after the responseUp to 15 min after the last requestWithin the same limitwaitUntil: 30 s. Alarms: 15 min
Durable workflowsNot yet, plannedWorkflow SDK, billed per eventCloudflare Workflows
Resumable streamsThe AI SDK’s resume, without RedisYou add Redis, or use WorkflowsBuilt in
Conversation storageTurso databases, includedPartners (Neon, Upstash, Supabase)SQLite in each agent
LanguagesJavaScript and TypeScript (Node.js, Bun)JavaScript, TypeScript, PythonJavaScript and TypeScript
WebSocketsNoBeta, closed at the function’s limitYes, the main transport
EU regionFrankfurt, alwaysFrankfurt and others, if chosen. Default: Washington, D.C.Agents’ data, by jurisdiction. Workers: Enterprise add-on

Run your agent at a price you know today.

We use your email only to invite you to the beta. Privacy