Storage

Bcachefs: The Best Filesystem Nobody Can Ship

A filesystem with the best ideas in Linux storage got merged into mainline, thrown out of mainline, and told to ship itself out-of-tree. The technical story and the politics story are both worth your attention.

Here’s a story that shouldn’t make sense.

In October 2023, bcachefs was merged into the Linux kernel. In June 2025, Linus Torvalds announced it was being thrown back out. By September 2025, the code had been deleted from the mainline tree entirely, and the project now ships as an out-of-tree DKMS module you install from a third-party apt repository.

The reason wasn’t technical. It wasn’t performance, and it wasn’t correctness. It was a breakdown in the review relationship between the maintainer and the project lead.

Which leaves the awkward question every infrastructure engineer eventually has to answer: the technical work is genuinely excellent, so should I run this in production?

I went through the design docs and the release history to work out where I land.

What it actually is

First, some orientation. bcachefs descends from bcache, which was a block-layer cache: the thing you slap in front of a slow disk to make it look fast. bcachefs took the ideas and built an actual filesystem out of them. Announced in 2015, developed by Kent Overstreet.

The interesting design decision shows up immediately, and it’s the one that makes everything else possible:

“The internal architecture is very different from most existing filesystems where the inode is central and many data structures hang off of the inode. Instead, bcachefs is architected more like a filesystem on top of a relational database, with tables for the different filesystem data types: extents, inodes, dirents, xattrs, et cetera.”

No central inode. Separate tables for extents, inodes, directory entries, and extended attributes, indexed by a hybrid B+ tree.

That matters more than it sounds. The classic Unix filesystem couples a file’s metadata to a single inode, which means recovery has to treat that inode as a special, critical structure, and concurrency has to be serialized around it. If your filesystem is table-oriented, you get much better recovery semantics, much easier snapshotting, and much better parallelism, because different parts of the tree don’t contend for one lock.

On top of that layout you get the standard modern feature set: copy-on-write, checksumming of both data and metadata, multi-device support, snapshots, reflinks, transparent encryption, and compression (lz4, zstd, gzip).

The part I think is genuinely novel

Erasure coding. Not RAID-as-a-block-layer: the filesystem itself does the coding.

bcachefs implements Reed-Solomon, the same algorithm as RAID5/6. So on the surface, nothing special. Where it diverges is in the two problems that have historically made filesystem-level erasure coding a bad trade.

The write hole. In conventional RAID, a small write that partially covers a stripe requires a read-modify-write of the whole stripe. If power drops between the read and the write, the stripe is inconsistent. This is the write hole, and it’s why RAID5/6 spent decades with battery-backed write caches and awkward flush semantics.

bcachefs avoids it structurally: no in-place updates on existing stripes. Per the project, erasure coding is “high performance with no write hole.”

Fragmentation. ZFS’s equivalent feature, when you enable erasure coding, fragments incoming writes: you have to buffer and align them into full stripes. bcachefs builds stripes asynchronously instead: incoming writes are initially replicated, and the extra replicas are dropped once the stripe is complete.

So you get erasure coding’s capacity benefit without ZFS’s write path penalty, and without a write hole to defend against. That’s a real contribution, not a marketing line.

Compression is per-inode rather than only per-filesystem, which is more flexible than most alternatives. And the “backpointers” mechanism (where the allocator can discover which other references point to a given physical block, from anywhere in the tree) is a clean solution to the shared-block problem that a table-oriented design makes possible.

How it got ejected

Now the part that’s genuinely strange to look at as an engineer, because you go looking for a technical smoking gun and there isn’t one.

Oct 2023: merged into Linux 6.7. Released January 2024.

Apr 2024: Torvalds, on 6.9-rc3, when told bcachefs was stable: “if you thought bcachefs was stable already, I have a bridge to sell you.” In August 2024, on stability expectations again: “nobody sane uses bcachefs and expects it to be stable.”

That August 2024 line is a marketing problem as much as a technical one. A filesystem nobody is expected to be able to rely on is very hard to get anyone to deploy. It’s also arguably unfair (plenty of code in mainline has had rough edges for a long time), but the expectation is the product.

Same month, the Debian maintainer orphaned bcachefs-tools in Debian.

Nov 2024: Overstreet was restricted by Linux’s Code of Conduct Committee from sending contributions during the 6.13 cycle, cited as “written abuse of another community member.”

Jan 2025: patches for 6.14 merged without issue.

June 2025: Torvalds refused a bcachefs pull request for 6.16-rc3 and said:

“I think we’ll be parting ways in the 6.17 merge window.

You made it very clear that I can’t even question any bug-fixes and I should just pull anything and everything.

Honestly, at that point, I don’t really feel comfortable being involved at all, and the only thing we both seemed to really fundamentally agree on in that discussion was ‘we’re done’.”

Aug 2025: marked EXTERNALLY MAINTAINED in MAINTAINERS. The code stayed in-tree deliberately, to smooth the transition.

Sept 2025: DKMS out-of-tree plan announced. Then, for 6.18, the code was deleted outright (commit f2c61db):

“bcachefs was marked ‘externally maintained’ in 6.17 but the code remained to make the transition smoother. It’s now a DKMS module, making the in-kernel code stale, so remove it to avoid any version confusion.”

Read that commit message again. “Remove it to avoid any version confusion.” The stated risk isn’t data loss or instability; it’s two versions existing and confusing people. That’s a housekeeping reason, which is almost strange for something that started with a breakdown in trust.

June 2026: bcachefs-tools 1.38.6, “the performance release.” The experimental label came off. The Reconcile operation (formerly “rebalance”) got faster and more parallel. Per the project notes, erasure coding is now “in use and seems to be working quite well.” Benchmarks from that cycle showed real-world gains, including faster compile times on large open-source projects, which is a nice sanity check that it’s not just synthetic.

So: the project is healthy, shipping steadily, and improving. It’s just permanently outside the kernel.

So should you run it?

Here’s my honest assessment, and it’s split.

For

  • The design is better than what you have. If you’re on ext4 with an mdraid/LVM stack underneath, you’ve assembled a filesystem and a redundancy layer from separate parts, and you get the RAID write hole anyway. bcachefs does integrity and redundancy in one coherent layer.
  • Erasure coding without the write hole is the headline. If you’re capacity-constrained and currently paying for RAID6 parity, this is the real prize.
  • Data checksumming by default. ext4 gives you metadata checksums only. bcachefs checksums data too, which catches silent corruption: bit rot, firmware bugs, failing disks lying about their health. In a healthcare or financial processing environment, silent corruption is a much bigger deal than throughput.
  • It moved. A project that survived ejection from mainline and is still shipping releases is not a hobby project.

Against

  • Out-of-tree is a real operational tax, and this is the one I’d weigh most heavily. A DKMS module means your filesystem depends on compiling against every kernel you install. Kernel update, then module build, then a chance the build fails or the ABI shifted. For a bare-metal box you reboot on a schedule, that’s manageable. For fleet-managed or compliance-locked environments, “the filesystem is not in the kernel” is often a disqualifying answer regardless of merit, because now you own a kernel module’s maintenance across your whole hardware refresh cycle.
  • The stability expectations have a history. The “nobody sane expects bcachefs to be stable” line is from 2024, and the “experimental” label only came off in June 2026. That’s a two-year gap of nobody expecting much, followed by a label removal. I’d want to see a few more release cycles before I’d bet on it.
  • Ecosystem maturity. Recovery tooling, documentation, and the volume of battle-tested operational knowledge are all thinner than btrfs. When it goes wrong at 2am, you’ll be reading source.

What I’d actually do

Don’t put it on anything you can’t afford to rebuild. That’s the honest position. It doesn’t mean “never.”

Do treat it as a great fit for specific, well-understood workloads where you can prove the benefit: high-capacity storage where RAID6 parity is costing you real money, or where data checksumming across a large array has genuine value. Not as a general-purpose root filesystem.

Re-evaluate in a year. This is the part most advice about bcachefs gets wrong by being frozen in 2024. The project is actively shipping and improving; the situation changes. The reasons to avoid it today are substantially about process, and process can change. If you’re reading this in a year and I’m out of date, that’s a fair hit.

Trying it

It’s distributed as a DKMS module, like ZFS, not packaged in Debian proper.

# Add the project's repository, then:
sudo apt update
sudo apt install bcachefs-tools

Single device, no redundancy (the minimum viable test):

sudo mkfs.bcachefs /dev/sdX1
sudo mount -t bcachefs /dev/sdX1 /mnt/test

Multi-device with replication (closer to how you’d actually deploy it:

sudo mkfs.bcachefs \
  --fs_size=8G \
  --data_replicas=2 \
  /dev/sdX1 /dev/sdY1 /dev/sdZ1

sudo mount -t bcachefs UUID=<your-uuid> /mnt/test

Options can be set at format time, at mount time, at runtime via /sys/fs/bcachefs/<fs>/options/, or per-inode through extended attributes. That last one is what lets you compress one directory and not another.

Do this in a VM first. A throwaway Debian instance with a couple of loopback files, or better, three small virtual disks, so you can actually pull a device out and watch what happens. Anything less and you’re just running a filesystem; you’re not testing the part that matters.


This one is an analysis rather than a report. I didn’t run bcachefs in production and I’m not claiming to have; I’m reading the design docs, the project notes, and the release history, and forming a view. If you have run it somewhere real, I’d like to hear how recovery went and how the DKMS upgrade cycle actually treated you; that’s the specific gap in my information.

Sources

Comments are reviewed before they appear. Yours will be published once it has been approved.

Leave a Reply

Your email address will not be published. Required fields are marked *

Share with