← BLOG

Engineering · Sep 30, 2026 · 11 min read

Should you build on Cloudflare? What we learned from D1 and Durable Objects

IR
Illiana Reed
CEO & Product, camelAI
Should you build on Cloudflare? What we learned from D1 and Durable Objects

In July, we wrote about why we moved camelCode, our open-source coding agent, into a Cloudflare Durable Object. In early September, we started replacing Cloudflare in our agent infrastructure with camelRun, a hosted runtime for AI agents that we built. It's open source on GitHub.

Durable Objects made the agent cheaper and faster than it was on VMs. But each Durable Object gets 128 MB of memory, which a long-running agent can outgrow. One long chat ran out of memory about 2,500 times in two days. On a separate, much simpler project, D1 said our 5.6 MB database was overloaded and rejected even a plain SELECT 1. In both cases, the error told us what failed but not why.

I'd still start almost any project on Cloudflare Workers, because they're the fastest way I know to ship something. But for anything with paying users, especially for the parts customers depend on most, like your main database or a long-running agent, I would not choose D1 or Durable Objects.

D1: overload errors on a 5.6 MB database

Alongside camelCode, we ran a much simpler inference project on Workers. It used a handful of Cloudflare products, all in the standard way the docs describe. Its shared D1 instance held the usual things, like accounts, API keys, and billing.

D1 is a good default for a new project. It's SQLite, so there's nothing to set up. A Worker talks to it through a binding, so there are no connection strings or pools to manage.

Our usage was light. The whole database was about 5.6 MB and 15,000 rows. We ran a few hundred queries a minute. Our two most common lookups, checking an API key and checking a login session, averaged under half a millisecond and read two rows each.

In early September, D1 rejected those lookups in bursts with D1 DB is overloaded. Requests queued for too long. Every API call checked its key first, so API calls failed too. That day we added indexes to the auth lookups, turned on read replication, and cached key checks in each account's Durable Object. The errors kept coming back. During the bursts, even a plain SELECT 1 failed. Even Cloudflare's own API for reading the database's info failed.

Searching the error turned up other people with the same problem on small, quiet databases. In April, an engineer on the D1 team explained the cause in a community thread. A database can land on a machine with a noisy neighbor, where expensive storage work for another database starves yours of CPU. The engineer said the team is changing the threading model for Durable Object storage so this can't happen, moved that user's database to a different machine in the meantime, and advised against retrying overload errors, because retries can make the overload worse. In June, Cloudflare support gave the same explanation to a user whose idle 40 MB database was throwing the same error we saw.

Until the fix ships, you can't control whether a busy neighbor takes your database down. I could accept a database getting slower when a neighbor is busy, or maintenance operations being unavailable for a while. No app can afford a shutdown of primary read/write operations.

We moved the project's shared database to PlanetScale Postgres, connected through Hyperdrive, and verified every row across all 24 tables before switching. The project stayed on Workers, and Durable Objects kept handling per-account usage and concurrency, which is the kind of small, fast work they're good at. Cloudflare supports this path well. Hyperdrive has a built-in PlanetScale integration, and as of June you can create PlanetScale databases right from the Cloudflare dashboard. The migration was very easy. GPT shipped the migration to prod while we were eating lunch, without our approval, and there were no bugs. No harm, no foul.

camelCode: Durable Objects and long-running agents

On camelCode, a single chat could touch Workers, several Durable Objects, R2, and containers. When something failed, it was hard to tell where an error started versus cascaded down from. We had assumed a lot of our issues came from our own complexity, but after months of debugging and seeing others hit the same problems, we think the problem is a fundamental mismatch between agents and Durable Objects.

Durable Objects are small, light, and fast, which makes them great for small, light work. Moving our agent into them cut our costs and our latency compared to VMs. But agents are large and complex, and they run for a long time. A coding agent can work on one task for tens of minutes, and agents like Codex and Claude Code can run for hours. In our experience, Durable Objects are comfortable with work measured in seconds, and minutes starts to push it. The longer an agent runs, the more likely something restarts it partway through.

Restarts

Every deploy restarts every running Durable Object. Cloudflare's docs say runtime restarts will also evict objects at unpredictable times. A stateless request recovers beautifully. An agent that has been running for 20 minutes has to restart. A single deploy can do this to a lot of users at once. This means every turn has to be resumable. It has to pick up where it left off, in a reasonable amount of time, without corrupting its state. We built in that complex recovery code. Our agent journaled its progress as it worked and resumed the turn after a restart. One production thread survived 15 restarts in a row and still finished correctly. But it took a lot of heavy code, each restart added latency, and ultimately your agent loop should not have to survive 15 restarts.

Memory

Memory is the other hard limit. Each Durable Object gets 128 MB to hold the conversation, tool calls, and everything else, like library imports and subagent outputs. When a Durable Object goes over, it's reset, and the caller gets Durable Object's isolate exceeded its memory limit and was reset. An agent in a Durable Object has to be ready to die at any moment.

In August, one very long chat thread crashed its Durable Object about 2,500 times over two days. Each time the object woke up, it loaded that thread's saved stream into memory, went over the limit, and was reset before our cleanup code could run. We fixed it by capping every buffer and adding a circuit breaker that stopped replaying a thread's stream after three crashes in a row. If a turn kept crashing anyway, each retry used less memory than the last, and as a last resort the agent saved what it had finished and asked the user to continue. A Cloudflare engineer reported the same failure in Cloudflare's Think harness in July, where one noisy pnpm install pushed the agent's Durable Object over its memory limit and left the session crashing on every wake.

Cost

Durable Objects are intended to be small. Small is also supposed to mean cheap, since you aren't paying for a whole VM. With heavy use, the costs added up faster than we expected. Duration is billed on the full 128 MB whether you use it or not, and SQLite rows read and written are billed on top. An agent that checkpoints so it can survive restarts writes a lot of rows. The recovery work showed up on the bill too.

Fast cold starts were the other benefit we expected. On a VM or sandbox provider, a cold boot can take a few seconds, so a user's first message waits unless you pre-warm. When we first moved our agent into Durable Objects, first messages started much faster, and on average they stayed that way. However, we saw logs of cold starts over 10 seconds, and there's no stated baseline to design around.

camelCode ran on Durable Objects for about 2 months, and it worked because we built around each of these limits. For a small team, it was a very high engineering tax that never worked smoothly, and it was a big part of why we replaced Cloudflare in our agent infrastructure.

Debugging

In my opinion, this is where Cloudflare has the most room to improve.

The first problem we couldn't explain was WebSockets. Users randomly could not connect their chat over a WebSocket. This meant paying users could not use the product at all. The best fix was a hard reset of the tab, but even that was not a guaranteed fix. The cause could have been Cloudflare, their router, or their internet provider, and we had no information to tell which. At the time, we fixed this by falling back to HTTP polling as soon as a socket errored or closed unexpectedly. If you're building on Cloudflare's WebSockets, I'd plan for a fallback.

When a Durable Object runs out of memory, the reset error doesn't say what was using the memory. Cloudflare supports heap snapshots in local development, but our crashes came from production data, like that one long thread, which is hard to reproduce locally. Without a view into production memory, you end up guessing.

For D1, all we got was the overload error, on queries that took under a millisecond. The only explanation we found was in a community thread.

Cloudflare has started on this, and in August they launched agent tracing, with session replay and a trace waterfall for every model and tool call. That's a good step. The gaps that hurt us were lower in the stack: why a connection failed, what was using memory before a crash, and why a database with almost no load was overloaded.

What we'd ask Cloudflare for

The first two are close to bug fixes:

  • Guarantees. We'd like some number of nines on D1 staying available when a neighbor is busy.
  • Memory visibility. Even with a bigger limit, agents will go over sometimes. When they do, you need to see what was using the memory.

The rest are about agents:

  • Bigger Durable Objects. Agents are large, and 128 MB has to hold your code, your libraries, and all of the agent's state.
  • Fewer surprise restarts. Objects shouldn't get killed out of nowhere. We'd like an easier way to hook into their lifecycle, like a chance to checkpoint before a deploy replaces them.
  • An Agents SDK that works well inside these limits. Cloudflare is clearly serious about agents. This year they shipped Project Think, a harness built on the Agents SDK, and in August they open-sourced Cloudflare OS, a workspace where teams build apps with agents. I suspect making those work well will mean fixing much of what's in this post, because these limits are as hard for Cloudflare's harness as they were for ours.

So, should you build on Cloudflare?

For most projects, I'd say yes, at least to start. Workers are fast to build on and cheap to run. Deploying is one command, rolling back is just as easy, and each version can get its own preview URL, so staging is simple. Give your agent a Cloudflare API token and it can create the database, deploy the app, and roll it back without you touching a dashboard.

Cloudflare comes with nearly every primitive a project needs: a database in D1, object storage in R2, Queues, durable Workflows, locks and coordination through Durable Objects, and inference through Workers AI. You could give your agent a VM instead, and it would get you most of the way. You'd still have to provision storage and inference endpoints yourself, and keep track of which bucket belongs to which app so nobody deletes it or changes its permissions. On Cloudflare, the app and everything it uses live together in one project.

Cloudflare can get you meaningfully far, especially for personal software or internal tools for your team. If something takes a few extra seconds one day, or throws an error, you can live with that. Once you have hundreds of users who depend on your product, those same errors show up in conversion and retention, and almost nothing is worth losing those. Workers and Durable Objects were built for fast, small pieces of work, and they're great at it. For the parts your customers depend on most, like your main database or a long-running agent, I'd use the boring, well-understood stack, which is why we moved our D1 database to Postgres.

Our agent was harder to move than the database. We built camelRun to fix the restart, memory, and debugging problems we ran into, and camelCode runs on it now. camelRun is a hosted runtime for AI agents. You define an agent's model, instructions, tools, and channels. camelRun runs the agent loop, stores its history and files, runs model-written code in a sandbox, and scales your agents. It's open source under the AGPL, and the code is at github.com/qaml-ai/run. We'll write more about how it works soon.

TL;DR

  • camelCode, our open-source coding agent, ran on Cloudflare Workers and Durable Objects for about 2 months. We still think Cloudflare is one of the best places to start a project.
  • On a separate project, D1 rejected sub-millisecond queries on a 5.6 MB database with D1 DB is overloaded. Requests queued for too long. Indexes, caching, and read replicas didn't stop it. The D1 team has described this as a noisy-neighbor problem they're working to fix. We moved that database to PlanetScale Postgres through Hyperdrive, and the migration was easy.
  • Durable Objects restart on every deploy and on runtime restarts, and reset when they go over 128 MB. We built journaling and recovery so agent turns could survive restarts, but it took a lot of code, and every restart added latency. One thread crashed about 2,500 times in two days before we capped its buffers.
  • Durable Objects were cheaper and faster than VMs on average, but checkpointing added to the bill, and some cold starts took over 10 seconds.
  • Debugging is the biggest gap. WebSocket failures, memory crashes, and D1 overloads came with little or no explanation. We worked around the WebSocket failures by falling back to HTTP polling.
  • We'd ask Cloudflare for D1 availability guarantees, production memory visibility, bigger Durable Objects, fewer surprise restarts, and an Agents SDK that works well inside these limits.
  • For prototypes, MVPs, and internal tools, Cloudflare is hard to beat. For anything with paying users, especially your main database or a long-running agent, we'd pick the boring, well-understood option.
  • We moved our agent infrastructure off Cloudflare and onto camelRun, a hosted runtime for AI agents that we built to fix these problems. It's open source under the AGPL at github.com/qaml-ai/run.

FROM THE TEAM BEHIND CAMELAI

Try what we're building.