Remix.run Logo
Show HN: EterDB, a Postgres fork that makes it easy to recover from incidents(eterdb.com)
33 points by fdeth 4 days ago | 15 comments

Hey everyone, a few months ago, after recovering from a Clade-generated bug, I was thinking to myself “wouldn’t it be nice if prod DB writes were easy to roll back”. I’ve put something together: it’s a two-container deploy plus a CLI tool. It’s far from being production ready, so please be gentle. All feedback welcome!

The nitty gritty: https://eterdb.com/tech

GitHub: https://github.com/eterdb/eterdb

micw 4 hours ago | parent | next [-]

The idea looks interesting. But when I think of "make incident recovery easy", the very last thing I want is a fork that differs from the standard that everyone else is running. I would have a better feeling if that would be an extension, not a fork.

maxloh 3 hours ago | parent [-]

From the README,

  It is PostgreSQL 18. One piece, read-dependency capture, has to be in the engine, and it ships as a small upstream-tracked patch. Extensions including pgvector, your ORM, and your SQL dialect all work unchanged.
https://github.com/eterdb/eterdb#how-it-works
0c3ca83 2 hours ago | parent [-]

Yes, so, in other words it's patching the internals. And the "how it works" is vibeslop.

hasyimibhar 2 hours ago | parent | prev | next [-]

The homepage shows an example of an agent accidentally dropping a table.

1. This is such an insane example, I don't get why would you give agents write access to your production db in the first place. Are people really doing this? I don't even have production connection urls on my laptop. Any manual statements executed against the db must be treated as a war-room situation with at least another engineer reviewing your SQL before you execute it.

2. Dropping the table could have easily caused writes to fail. Most likely there is no way to recover these writes (especially if it's from user requests), so it could have lead to loss of data. Reversing the table drop doesn't fix this issue.

traceroute66 13 minutes ago | parent [-]

> Reversing the table drop doesn't fix this issue.

Also reversing a table drop doesn't solve extra corruption routes such as cross-table dependencies.

dewey an hour ago | parent | prev | next [-]

So instead of doing the easy thing (Giving your Claude a read-only role, having snapshots etc.) you decided to patch the database which needs to be kept in sync with every PG release and build a landing page?

traceroute66 2 hours ago | parent | prev | next [-]

> after recovering from a Clade-generated bug, I was thinking to myself “wouldn’t it be nice if prod DB writes were easy to roll back”.

Hmmmmmm......

1. "Claude-generated bug". No it was PBCAK (Problem Between Chair And Keyboard) a.k.a "foolish person ran Claude against the production database without testing it elsewhere". There, fixed it for you.

2. This "product" is solving a problem that is already solved. You can for example use a SaaS provider such as Aiven[1] who will provide you with PITR (Point-In-Time Recovery) point and click solutions. Alternatively there is more than one piece of Postgres backup software that lets you do the same on a DIY basis.

3. "Out of the box" you have pg_dump. You could have just done a simple pg_dump before letting Claude loose on your database.

[1] https://aiven.io/

MarceColl an hour ago | parent [-]

I have no relation to the project, but calling this problem solved and then offering a vastly inferior solution to the proposal (PITR) is a bit unfair

traceroute66 15 minutes ago | parent [-]

Point one remains. The "solution" appears to come from somebody who thinks its OK to run Claude against a production database.... “wouldn’t it be nice if prod DB writes were easy to roll back”

The point remains that if they ran Claude against a test database they would have found the bug without killing their production database, and therefore also not need to come up with an over-engineered "solution".

Sometimes also the less over-engineered the better. Stuff like PITR and pg_dump is battle-tested and easy to reason about.

themgt an hour ago | parent | prev | next [-]

Forking Postgres and patching the transaction engine to aid granular recovery is one obvious way to achieve this, but the more robust solution is to fork the Linux kernel and apply a surgical patch to the TCP/IP stack. A kernel level regex check on port 5432 traffic prevents Claude from dropping your prod DB tables in the first place.

Haven880 41 minutes ago | parent [-]

Wouldn't it be easier to patch gnu c compiler that compiles the kernel? That way anyone compiling kernels will auto include the enhancement by default across all OS, unix or linux.

NewJazz 3 days ago | parent | prev | next [-]

Uh, pitr is a thing in many many postgresql management layers.

https://pgbarman.org/

justinclift 14 minutes ago | parent | next [-]

I think the think they're aiming for here isn't to rewind to a particular point in time, but to remove a transaction from the history and while keeping the transactions after it.

fdeth 3 days ago | parent | prev [-]

Yes, but PITR is not surgical, you can’t easily undo just the faulty transactions.

thisismyopinion an hour ago | parent | prev [-]

Slop.