rubbie kelvin.
software engineer, nigeria.
sysconf · 2026
the autopsy of an asynchronous deadlock.
cpu
42%
/health
200 OK
checked every 1s
errors
0
last 5 min
orders completed, last 5 min
1,284
last order just now
nothing.
first, the shape
a deadlock is a loop of waiting.
task a waits for task b. task b waits for task a.
nobody is broken. nobody will ever move.
async, for everyone
await means: wake me when it's ready.
| language | the same idea |
|---|---|
| javascript | const res = await fetch(url); |
| python | res = await session.get(url) |
| c# | var res = await client.GetAsync(url); |
| rust | let res = client.get(url).send().await?; |
a few threads run thousands of tasks.
a task that awaits gives its thread back, and waits on a shelf until something wakes it up.
why it's hard
a stuck task isn't on any stack.
a stuck thread
the debugger finds it parked on a lock
its stack names the exact line
five minutes, and you're done
a stuck task
every thread is idle, waiting for work
none of your code is on any stack
the task is just data, waiting for a wake-up that never comes
and often, half the loop isn't in your process at all. it's in your database.
scenario 1
the stockroom key.
you lock the stockroom and put the key in your pocket.
you send your assistant to fetch a box from it,
and wait at your desk for the box.
your assistant waits at the locked door. for you.
you wait for them. they wait for you. nobody is busy. nobody complains.
scenario 1 · the stockroom key
the code.
async fn mark_paid(
db: &mut Client,
id: i64,
) -> anyhow::Result<()> {
let tx = db.transaction().await?;
// locks row 42
tx.execute(
"update orders set status = 'paid' where id = $1",
&[&id],
).await?;
// waits for the worker…
send_receipt(id).await?;
// never reached
return Ok(tx.commit().await?);
}async fn send_receipt(id: i64) -> anyhow::Result<()> {
// a different connection
let db = POOL.get().await?;
// waits for row 42
db.execute(
"update orders set receipt_sent = true where id = $1",
&[&id],
).await?;
return mailer::send(id).await;
}the stockroom
row 42
the key in your pocket
the open transaction
your assistant
send_receipt
waiting at your desk
.await
scenario 1 · the stockroom key
why nobody catches it.
the handler awaits the worker's reply.
the worker waits for the lock on row 42.
row 42's lock is held by the handler's open transaction. the loop is closed.
what the database sees
a queue, not a loop: the worker waits behind a session that's doing nothing. its deadlock detector stays quiet.
what the runtime sees
two idle tasks. cpu 0%. health check green.
what only your code knows
the handler awaits the worker. that arrow closes the loop, and no tool sees it.
put together
the loop closes through an await. only you can see the whole thing.
scenario 1 · the autopsy
step 1: ask the database who's waiting.
select pid, application_name, state, wait_event,pg_blocking_pids(pid) as blocked_byfrom pg_stat_activity;
| pid | application_name | state | wait_event | blocked_by |
|---|---|---|---|---|
| 812 | receipt-worker | active | transactionid | {804} |
| 804 | orders-api | idle in transaction | — | {} |
the lock holder is us, and it's idle. doing nothing is the bug.
illustrative output · mysql: sys.innodb_lock_waits · sql server: blocking_session_id in sys.dm_exec_requests
scenario 1 · the autopsy
step 2: ask the runtime.
tokio-console, tasks view (illustrative)
| task | state | idle for | polls |
|---|---|---|---|
| mark_paid | idle | 4m 12s | 3 |
| send_receipt | idle | 4m 11s | 2 |
| heartbeat | idle | 1s | 253 |
the same view elsewhere: python asyncio.all_tasks() · c# parallel stacks › tasks · go goroutine dump
the database said
the worker waits for the handler's transaction.
the runtime says
the handler waits for the worker.
put them together
a loop.
neither tool shows a deadlock. you draw the last arrow yourself.
scenario 1 · the fix
give the key back before you send anyone.
// orders.rs - mark_paid(...)
let tx = db.transaction().await?;
tx.execute("update orders …", &[&id]).await?;
send_receipt(id).await?; // stuck
tx.commit().await?;// orders.rs - mark_paid(...)
let tx = db.transaction().await?;
tx.execute("update orders …", &[&id]).await?;
tx.commit().await?; // lock released
send_receipt(id).await?;the rule
never await anything while a transaction is open.
scenario 2
half the hallway.
two people meet in a narrow hallway. each waits for the other to step aside.
a guard watching the camera sees it, and sends one of them back.
now the hallway turns a corner. one person stands just out of view.
the guard sees one person waiting. that looks normal.
scenario 2 · half the hallway
first, the one the database catches.
update accounts … id = 1;← locks row 1update accounts … id = 2;← waits for b
update accounts … id = 2;← locks row 2update accounts … id = 1;← waits for a
ERROR: deadlock detectedDETAIL: Process 130 waits for ShareLock on transaction 741; blocked by process 129.Process 129 waits for ShareLock on transaction 740; blocked by process 130.
both locks are in the database, so the guard sees the whole loop. about a second later, one side gets an error it can retry.
scenario 2 · half the hallway
now move one lock into your app.
async fn reserve(app: &App, id: i64) -> anyhow::Result<()> {
// app lock (a tokio::sync::Mutex)
let mut pending = app.reservations.lock().await;
// waits for restock
app.db.execute(
"update products set stock = stock - 1 where id = $1",
&[&id],
).await?;
pending.push(id);
return Ok(());
}async fn restock(app: &App, id: i64) -> anyhow::Result<()> {
let mut db = app.pool.get().await?;
let tx = db.transaction().await?;
// row lock
tx.execute(
"update products set stock = stock + 5 where id = $1",
&[&id],
).await?;
// waits for reserve
let mut pending = app.reservations.lock().await;
pending.retain(|&r| r != id);
return Ok(tx.commit().await?);
}scenario 2 · the autopsy
four arrows, two witnesses.
reserve waits for row 7.
row 7 is held by restock's open transaction.
restock waits for the app lock.
the app lock is held by reserve. the loop is closed.
what the database sees
reserve waiting on an idle-in-transaction session from your own app: the same picture as scenario 1. the rest of the loop is in your code.
what the runtime sees
restock is waiting for the app lock. runtime tools show who waits for a lock, never who holds it.
what only your logs show
who took the app lock, and when. that's the last arrow.
12:03:07.114 reserve: took the reservations lock (product 7)
put together
two arrows from the database, one from the runtime, one from your logs. draw them all and the loop closes.
scenario 2 · the fix
bring the lock back on camera.
best
one source of truth. keep the reservation in the database, so both locks are on camera and a mistake becomes a retryable error within a second.
or
don't hold an app lock across a database call. do the database work first, then lock, update memory, release. no await in between.
guardrail
lock_timeout, so a stuck wait becomes an error instead of a freeze.
the method
the autopsy, in any stack.
01
is it a deadlock?
cpu near zero, progress flat, heartbeat still ticking. waiting longer changes nothing.
02
list every waiter
for each stuck thing: what does it hold? what is it waiting for?
03
draw the loop
follow the arrows until you're back where you started. note which tool saw each one.
04
break one arrow
release before you await. one lock order. a deadline on every wait.
the method
where to look.
| layer | the question | tools |
|---|---|---|
| database | who waits for whom? | postgres: pg_stat_activity, pg_blocking_pids() · mysql: sys.innodb_lock_waits |
| runtime | which tasks stopped being woken? | rust: tokio-console · python: asyncio.all_tasks() · c#: parallel stacks › tasks · go: goroutine dump |
| your code | who took which lock, and when? | logs around every lock · timeouts that report instead of cancel |
| process | is a thread stuck in our code? | lldb / gdb, or your platform's thread dump |
a timeout that reports
logs "still waiting after 5s" and keeps waiting, so the evidence stays where it is. node.js has no built-in task inspector, so your own logs do this job.
before the next one
five habits.
- give every service a heartbeat and a progress counter.
- log when you take a lock, not only when you wait for one.
- use timeouts that report, not cancel.
- set idle_in_transaction_session_timeout and lock_timeout.
- never await while holding a transaction, a lock or a connection.
thank you.
questions?