SONiC: A Network OS Built Like a Container Platform
Everyone else shipped a monolithic network OS. SONiC took a Debian box, put each function in its own container, wired them together with Redis, and called it a network operating system. It's the most interesting architectural bet in networking, and it's actually shipping.
Most network operating systems are a single enormous binary with a few helper processes bolted on. Cisco IOS, Junos, NX-OS, Arista EOS: all variations on “one big program owns the box.”
SONiC looked at that and did the opposite. It took a Debian host, put each functional domain in its own Docker container, wired them together with a Redis key-value database, and called the result a network operating system.
I’ve been following this project for a while, and I keep coming back to how much it changed about the conversation. Not just “white boxes are cheaper”; that was always true. It changed what people believed was possible for a NOS: a real, feature-complete L3 switch OS that Microsoft ran in production in Azure at enormous scale, contributed to by a genuine multi-vendor community, and that Cisco now ships on its own 8000-series routers.
Let me walk through how it actually works, because the architecture is the interesting part.
A note on the repos
Worth clearing up, since it trips people up: there isn’t one SONiC repo.
sonic-net/SONiC: the core platform. Where the containers, the build system, andswsslive. This is a Linux Foundation project as of 2022.aristanetworks/sonic: Arista’s platform support package. This is the one you linked, and it’s the right place to start if you want to see how a specific vendor’s hardware plugs in.Azure/sonic-buildimage: the build system that produces the images.Azure/sonic-platform-common: the base classes for thesonic_platformAPI that every vendor implements.
That split is the whole idea in miniature. The core OS doesn’t know what a Tomahawk is.
The architecture
SONiC is a NOS running on Debian Linux. Network functions are implemented as separate Docker containers, and they coordinate through a Redis instance.
The Redis part deserves a moment, because it’s doing more than you’d expect. It’s simultaneously:
- the configuration store (
CONFIG_DB): your intended state - the application state store (
APPL_DB): routes, next hops, neighbors, interfaces as they actually are - the hardware state store (
STATE_DB): port status, link state - the ASIC store (
ASIC_DB): the bridge to silicon - an IPC bus between processes that don’t share a memory space
So when you type a command, something genuinely interesting happens.
Following a single interface configuration
Say you configure an interface with an IP. Here’s roughly the path:
- Your CLI writes intent into
CONFIG_DB. - Inside the
swsscontainer,orchagent, the orchestrator, reads that intent out ofCONFIG_DB. orchagenttranslates it into a hardware operation and pushes it down tosyncd, viaASIC_DB.syncdis the container that actually talks to the ASIC. It loads a vendor-specific SAI implementation and calls into it.- The ASIC programs the hardware. State comes back up the same path into
STATE_DB.
Meanwhile, entirely separately, portsyncd is listening on netlink events (RTNLGRP_LINK) and synchronizing port status (speed, lanes, MTU, link state) into STATE_DB. intfsyncd does the same for interface events. neighsyncd handles neighbor resolution.
The pattern is consistent, and it’s why the naming is slightly confusing at first: everything ending in syncd is a synchronizer. Something in the kernel or the hardware happened; this process notices and syncs it to Redis. syncd is the odd one out; it synchronizes the software stack with the silicon.
Other containers do their own jobs: bgp (FRR/Zebra for routing), lldp, snmp, dhcp_server, dhcp_relay, teamd (LAG), nat, macsec, pfcwd (lossless priority flow control), ntp, pmon (platform monitoring), subintf.
The thing to internalize: no control-plane process owns the box. They all read intent from one place and publish state to another. That’s the entire architectural thesis, and everything else (the multi-vendor story, the testability, the operational model) falls out of it.
SAI is the actual unlock
Containers and Redis are interesting, but neither of those gives you multi-vendor. That comes from SAI (the Switch Abstraction Interface), specified through the Open Compute Project.
SAI is a vendor-neutral API for programming switching ASICs. Above the SAI line, SONiC is identical everywhere. Below it sits a vendor adapter that knows how Broadcom’s SDK, NVIDIA’s, or Cisco Silicon One’s SDK actually works.
This is the piece that makes the following sentence true: the same SONiC build runs on Broadcom Trident, Broadcom Tomahawk, NVIDIA/Mellanox, Intel, and Marvell silicon without changing the software above the abstraction.
Traditional NOSes don’t have this, and can’t. IOS runs on Cisco hardware because it’s a single codebase written against a single vendor’s SDKs. EOS is famously good engineering, but it’s Arista’s. The coupling isn’t cultural; it’s architectural.
You can see this directly in the Arista platform repo, where the ASIC support is a list of silicon families:
arista/components/asic/xgs/ tomahawk, tomahawk2..6, trident2..4
arista/components/asic/dnx/ jericho2, jericho2c, jericho3, ramon
arista/components/asic/aspeed/ ast2720
That’s a broad, current-generation silicon list (Tomahawk 5 and 6, Trident 4, Jericho 3) behind one interface. And Arista’s supported platform list runs from SOHO-ish CCS-720DT-48S through the 7050/7060/7170/7260/7280 families.
What vendors actually have to write
This is the part I think underappreciates the project. To bring up a new platform, a vendor implements the sonic_platform API (base classes in Azure/sonic-platform-common) and ships a package. Arista’s is four Debian packages:
sonic-platform-arista # system configuration files
sonic-platform-arista-libs # shared libraries
drivers-sonic-platform-arista # kernel modules and drivers
python3-sonic-platform-arista # the sonic_platform library
At boot, systemd services run an arista entry point that detects the platform, loads the right drivers, and then exposes fans, LEDs, transceivers, and sensors through sysfs. That’s the whole porting surface. A new box is a sonic_platform implementation plus an SAI adapter, not an operating system.
The part that surprised me: warm reboot
Here’s the question I had when I first looked at this seriously.
If state lives in Redis and processes are independently restartable, what happens when one of them dies? And what happens during a software upgrade: do you drop traffic?
The answer is warm reboot (swss warm-reboot). Because state is externalized to Redis rather than held in process memory, SONiC can reconcile the new software’s desired state against the hardware’s actual state, and apply only the difference.
So:
- No traffic drop on upgrade, because the forwarding state is preserved and only deltas are applied.
- A crashed container can be restarted without taking forwarding down, because the hardware keeps forwarding independently of whether the control plane is healthy.
- The Redis databases act as a durable record of intent, which makes a cold restart fast.
This is a direct architectural payoff from the “no process owns the box” design. You don’t get warm reboot for free. You get it because you externalized state on purpose.
I’d genuinely like to know how this behaves under real-world failure injection; that’s the specific thing I can’t learn from documentation.
Where SONiC stands in 2026
The question that used to be “is this a real project or a Microsoft science experiment” has a settled answer. It shipped.
Cisco runs SONiC on the 8000 series. This is the one that surprised me. Cisco, the vendor with the most to lose from open NOS adoption, ships a Cisco-validated SONiC for its 8000-series routers. Recent releases have added static weighted ECMP, VxLAN MAC rewrite, MAC fault error reporting, IP-in-IP tunnel termination config, SRv6 uN shift-forward support, online diagnostics, Teamd enhancements, and DHCPv4/v6 relay.
NVIDIA BlueField-4 is a 64-core Grace CPU with 800Gb/s networking, and SONiC is part of the “scale-in” story: using DPUs for offload next to the AI fabric. The DASH project (Disaggregated APIs for SONiC Hosts) targets exactly this: standardizing stateful offload services on smartNICs and DPUs while staying aligned with the SONiC ecosystem.
Enterprise SONiC 4.6 shipped with 13 new platform entries, including Celestica DS4100 (Tomahawk 4, 16×800G) and UfiSpace S9321-64EO (Tomahawk 5, 64×800G), plus safer routing and fabric performance work. The community platforms list keeps widening.
The commercial layer around it has matured too. The 650 Group has projected $8 billion in SONiC-related data-center switching revenue by 2027.
And Microsoft Azure has been running it in production at hyperscale for the better part of a decade.
The honest part
I’ve been enthusiastic, and it is a genuinely impressive piece of engineering. But there are things you should know before forming an opinion, and the community is better for being honest about them.
Production SONiC is rarely a pure-SONiC fabric. This is the caveat I’d lead with, and it’s the one most commentary skips. Real deployments tend to be mixed: SONiC on some roles, Cisco NX-OS or Arista EOS on others, in the same fabric. That’s completely legitimate engineering, and it tells you something real: teams adopt SONiC where its economics and features win, and don’t feel pressure to convert everything.
The operational model is a genuine shift, not a free upgrade. If your team’s muscle memory is “CLI → verify,” you’re now operating a distributed system. When something doesn’t work, the question becomes “which container, and what’s in which Redis database?” That’s a different debugging discipline. It’s a better architecture in a lot of ways and a steeper learning curve in all of them.
Support model is the real procurement question. You can run community SONiC yourself, or get a commercially supported distribution. Know which one you’re buying before you start, because it shapes everything downstream.
It won’t replace EOS or NX-OS broadly. Vendors have decades of installed base, mature tooling, and certified fabrics. SONiC is winning on economics, openness, and workload fit. That’s a big deal. It’s not a wipeout, and anyone telling you otherwise is selling something.
If you want to try it
Get a vs-style virtual SONiC image rather than touching hardware. The standard path is the VS build (sonic-vs.img.gz), which runs under QEMU and gives you a complete NOS with syncd backed by a virtual switch. You can build BGP, inspect the Redis databases, restart containers, and watch warm-reboot behavior without owning a switch.
Once you’re past that, the Arista platform repo is a genuinely great read; it’s small, it’s well-organized by component (chassis, fabric, linecard, platform, pfpu/xcvr), and it shows you exactly what “porting to new hardware” looks like in practice. Far more useful than another architecture diagram.
Things worth looking at once you have a running image:
# What's actually in the config store?
redis-cli -n 4 keys "PORT|*" | head
# How's the hardware doing?
redis-cli -n 6 keys "PORT_TABLE|Ethernet*"
# What containers are up?
docker ps
# Is orchagent healthy?
sudo container start swss
sudo sonic-cmd ...
Reading the actual contents of CONFIG_DB and APPL_DB is the single most clarifying thing you can do. It turns the architecture from a diagram into a thing you can watch move.
If you’ve run SONiC in production, particularly on Arista hardware or the VS images, I’d like to hear what surprised you. The failure modes and the operational rough edges are the parts I genuinely can’t get from documentation, and they’re the parts that would make this a better post.
Sources
- Arista Networks: aristanetworks/sonic, Arista platform support for SONiC
- sonic-net: SONiC and the Architecture wiki
- Cisco: SONiC Architecture and Cisco 8000 Series powered by SONiC
- Cisco: SONiC on Cisco 8000 release notes, 202511.1.1.0
- Sonic Foundation: sonicfoundation.dev
- NVIDIA: BlueField-4 scale-in network infrastructure
- r12f: Getting Started with SONiC, including the syncd/SAI deep dive
Comments are reviewed before they appear. Yours will be published once it has been approved.