About
I’m Zach. I’ve spent twenty-six years keeping other people’s systems running, mostly at three in the morning, usually for organizations that couldn’t afford to be down.
What I do
Infrastructure and network engineering, with a strong bias toward the parts of the job that don’t fit on a resume: the 2 a.m. page, the failed migration that has to be rolled back before the business opens, the storage array that started returning bad blocks and nobody had checked the logs in nine months.
The through-line is mission-critical operations: environments where uptime isn’t a KPI, it’s a precondition. I’ve worked across:
- Government and public sector: constrained budgets, long procurement cycles, and compliance requirements that don’t negotiate.
- Healthcare: where a slow system is a clinical safety issue, not just a support ticket.
- Financial services and credit-card processing: where seconds of latency are real money and a failed transaction is a reportable event.
- Education, nonprofit, and manufacturing: fewer people, wider responsibility, and a lot more systems than budget.
The variety taught me something specific: the hard problem is almost never the technology. It’s the constraint, the legacy dependency nobody documented, the vendor who won’t take a phone call, or the stakeholder who wants Tuesday’s change frozen until Thursday’s audit. Technology problems have answers. Organizational problems need negotiation, and that’s the part of the job most people underestimate.
How I work
Fail closed. If I can’t tell why something is safe, I don’t assume it’s safe. Ambiguity gets resolved before it gets deployed, not discovered during an incident.
Verify, don’t trust. I don’t believe a system is working because a dashboard is green, a vendor says so, or it worked yesterday. I pull the actual state and check it. This is also why I take apart software: I want to know what’s actually running, not what the documentation claims.
Document as you go, not after. The value of a runbook is entirely in whether it’s accurate when it’s needed under pressure, which means it gets written while you’re doing the work.
Automate the second time. Not the first. The first time you don’t yet know what the failure modes are. The second time, you do.
What I write about
This site is where I think out loud about infrastructure, networking, storage, and Linux. Some posts are things I’ve learned the hard way. Some are things I got wrong and had to fix. Some are just “this is genuinely interesting and I want to write it down.”
I’m not going to pretend everything here came from a production incident, and I’ll say so when a post is analysis rather than field experience. That distinction matters more than it might seem: the fastest way to stop trusting an engineer’s writing is to find out it was invented.
Currently
Building this site, which is a deliberate exercise in writing rather than just storing knowledge somewhere nobody will read.
If something here is useful, or you disagree with it, I’d like to hear both.
Get in touch
The most reliable way is LinkedIn. I’m always happy to talk shop about infrastructure problems, especially the kind where the obvious answer isn’t the right one.