The symptom
Production started throwing this on almost every endpoint:
PrismaClientKnownRequestError:
Invalid `prisma.jobPost.findMany()` invocation:
code: 'ETIMEDOUT',
meta: { modelName: 'JobPost' }
It hit every model (JobPost, Category, Session, CompanyProfile), including Better Auth’s session lookup, so users saw 500s. But it was not total: some requests worked and some failed, in bursts.
The stack stack: a Node 22 API (Hono + oRPC + Better Auth + Prisma 7 with the pg driver adapter) on Railway in Southeast Asia, talking to Neon Postgres in Sydney.
What I checked first (and ruled out)
Every one of these was a reasonable guess, and every one was wrong. Ruling them out quickly is the useful part.
| Suspect | How I checked | Result |
|---|---|---|
| Neon compute suspended (scale-to-zero) | Neon Monitoring graph | Active all day, no inactive band |
| IP allowlist blocking Railway | Neon project settings | ”IP restrictions: None set” |
| Database overloaded | CPU/RAM graphs, usage | Flat, near idle |
| Not using the pooler | Connection string | Already -pooler |
| Pool misconfiguration | Added connectionTimeoutMillis, idleTimeoutMillis, keepAlive, an error handler | Helped hygiene, did not fix it |
| Region too far apart | Sydney vs Singapore, about 100ms | Not far enough to explain it alone |
Two clues nagged at me:
- Failures happened exactly about 20 seconds apart in the logs.
- Prisma flattens the underlying driver error to just
code: 'ETIMEDOUT', hiding which address failed and why.
The test that found it
Prisma hides the real error, so I bypassed it. From the Railway service’s Console tab (a shell inside the running container), I opened 20 raw pg connections and printed the full error, including the per-address breakdown that Node attaches to an AggregateError:
cd /app/packages/database && node -e '
const {Pool}=require("pg");
(async()=>{
for(let i=0;i<20;i++){
const p=new Pool({connectionString:process.env.DATABASE_URL,connectionTimeoutMillis:8000,max:1});
const t=Date.now();
try{await p.query("select 1");console.log(i,"ok",Date.now()-t,"ms")}
catch(e){console.log(i,"FAIL",e.code,e.address,e.port,Date.now()-t,"ms",e.errors?e.errors.map(x=>x.address+":"+x.code).join(","):"")}
await p.end().catch(()=>{});
}
})()'
About half the connections failed, and every failure looked like this:
FAIL ETIMEDOUT ... 754 ms
3.106.69.51:ETIMEDOUT,
2406:da1c:b7c:d11:...:ENETUNREACH,
54.153.234.207:ETIMEDOUT,
2406:da1c:b7c:d20:...:ENETUNREACH,
54.206.85.193:ETIMEDOUT,
2406:da1c:b7c:d0e:...:ENETUNREACH
Two things jumped out:
- Every failure took about 753ms. That is suspiciously close to 3 × 250ms.
- Three IPv4 addresses timed out, three IPv6 addresses were unreachable. Six addresses, six attempts.
The actual cause
Since Node 20, outgoing TCP connections use autoSelectFamily (the “Happy Eyeballs” algorithm). When a hostname resolves to several addresses, Node tries them one at a time, and each attempt gets a default timeout of 250ms before it moves to the next.
Neon’s hostname resolves to three IPv4 and three IPv6 addresses. Here is what happened on each new connection:
- The three IPv6 addresses failed instantly with
ENETUNREACH, because Railway containers have no IPv6 route. - The three IPv4 addresses each got 250ms.
- From Railway in Singapore to Neon in Sydney, the round trip is around 100ms, and the TCP connect sometimes takes longer than 250ms once you add jitter.
- When all attempts hit the limit, Node threw
AggregateErrorwithETIMEDOUT, after about 750ms. - Prisma reported that as
code: 'ETIMEDOUT'with no other detail.
That is why it was intermittent. A connection that finished within 250ms worked, and one that did not, failed. It is also why it looked like a database outage: the error says “timed out”, and the message never mentions the 250ms limit.
The fix
Give each address more time and skip IPv6, using Node’s own flags. In Railway, add this environment variable to the API service (and every other service that connects to the database), then redeploy:
NODE_OPTIONS=--network-family-autoselection-attempt-timeout=5000 --dns-result-order=ipv4first
--network-family-autoselection-attempt-timeout=5000raises the per-address timeout from 250ms to 5s.--dns-result-order=ipv4firsttries IPv4 first, so the dead IPv6 addresses do not waste attempts.
I verified it before relying on it by re-running the same script with the variable set inline:
NODE_OPTIONS="--network-family-autoselection-attempt-timeout=5000 --dns-result-order=ipv4first" node -e '...same script...'
Before: about half the connections failed. After: 20 out of 20 succeeded.
You can also set it in code, before creating the pool. This keeps the fix next to the code that needs it:
import { setDefaultAutoSelectFamilyAttemptTimeout } from 'node:net';
// Node 20+ gives each address only 250ms to connect. Across regions
// (Railway SEA -> Neon Sydney) that is too short and shows up as ETIMEDOUT.
setDefaultAutoSelectFamilyAttemptTimeout(5_000);
A second, smaller problem: connections are expensive here
Even after the fix, each new connection took 1.8 to 2.5 seconds. Cross-region latency, TLS with certificate verification and authentication add up. That makes pool reuse matter. I had set idleTimeoutMillis to 10 seconds, which closes idle sockets quickly and forces expensive reconnects. If you use Neon’s pooler, raise it to around 60 seconds, since the pooler already manages idle connections on its side.
Prevention checklist
Use this before you ship any service that connects to a database in a different region or network.
- Put compute and database in the same region when you can. The best fix is no cross-region hop at all. If they cannot match, expect slower connections and plan for it.
- Set
NODE_OPTIONSfrom day one on any Node 20+ service that connects across regions or to a multi-IP hostname:--network-family-autoselection-attempt-timeout=5000 --dns-result-order=ipv4first - Run a connection soak test after every first deploy. Twenty raw connections from inside the container takes a minute and would have shown this immediately. Anything below 20/20 is a bug, not a flake.
- Never trust a flattened ORM error. When Prisma says
ETIMEDOUT, reproduce with the bare driver from inside the same container and printe.errors,e.addressande.port. - Look at the durations of failures. Failures that all take the same time (753ms, 20s) point to a fixed timeout somewhere, not random network loss. Divide by likely constants (250ms attempts, 3 addresses).
- Test from the real environment. My laptop connected fine. Only a shell inside the deployed container showed the failure, because the difference was Railway’s network (no IPv6), not the code.
- Set explicit pool timeouts (
connectionTimeoutMillis,idleTimeoutMillis,keepAlive) and attach apool.on('error', ...)handler, so a dead idle socket is logged rather than silent. - Use the provider’s pooled connection string (Neon
-pooler) for app traffic, and the direct one only for migrations. - Apply the variable to every service that talks to the database: API, workers and any app with its own Prisma client.
- Rule out the cheap suspects in order, and write down the result: is the database awake, are there IP restrictions, is it overloaded, is the URL correct? Do this before touching code.
Takeaways
ETIMEDOUTfrom a driver does not mean the server is down. It can mean your client gave up early.- Node 20’s 250ms per-address default is fine on a fast local network and quietly wrong for cross-region connections to hosts with several addresses.
- The most useful debugging step was removing the abstraction: a 10-line script with the raw driver, run in the real environment, printed the answer that Prisma had hidden.