It Worked Yesterday: Why Networks Fail at Scale
Tuesday, September 29 · 9:35–10:00 AM America/Denver · Main Room
Higher education networks increasingly resemble small cities—supporting tens of thousands of users, diverse device ecosystems, and mission-critical services spanning academics, research, healthcare, and student life. Yet many network failures in these environments are not caused by complex architectural flaws, but by accumulated technical debt and the realities of reactive operations.
This session presents real-world incidents from operating a 100G campus network, including wireless instability following a switch replacement, a statewide outage caused by a template misconfiguration, DHCP failures triggered by security controls, and a campus-wide disruption from a Layer 2 loop introduced by a third-party system. Each example demonstrates how seemingly minor decisions—carried-forward configurations, misunderstood scope, and hidden dependencies—compound over time and only surface under real-world conditions.
Attendees will explore how technical debt manifests in higher education environments, why many issues remain hidden until systems are stressed, and how organizational constraints often reinforce reactive approaches to network management. The talk will focus on practical lessons learned, including strategies to identify.
Zeb Whitehead
Auburn
“That doesn’t look right.” Zeb Whitehead has built a career on those five words. His badge says Network Architect at Auburn University, but “Network Wrangler” is closer to the truth — routing, switching, wireless, security, automation, and observability all fall under his watch, and he’s the guy people call when something is definitely not supposed to be doing that. He started as a Linux sysadmin, cut his teeth at Auburn University in the College of Engineering, and worked his way into running infrastructure at university scale. Off the clock, he’s on the farm — same instincts, different fence lines.