Network Engineer
Austin, TX
Role snapshot
You support the networks inside hardware validation labs around the world: the datacenter fabrics, timing, optical, and provisioning systems that let accelerator teams validate new platforms from first bring-up through production handoff. L2/L3 lab fabric design, precision timing as a service, intent based provisioning and automation, optical plus test infrastructure, lab operations and reliability, and cost takeout through decommissioning.
Required: network engineering
L2/L3 at depth: VLANs, loop prevention, LAG, IPv4/IPv6 addressing, SLAAC and router advertisements, DHCP relay and helpers, PXE boot flows end to end.
BGP: eBGP peering design, prefix policy, uplink preference tuning, per-rack or per-site route origination toward the border and aggregation layer.
Datacenter fabrics: leaf/spine role design, EVPN/VXLAN overlays, migrating racks from L2 bridged to L3 routed boundaries.
IPv6 operations: first-hop security (RA guard, address locking), DHCPv6 relay and MAC-stable allocation behavior including client identifier pitfalls, /64 per rack scale planning and fabric design
Precision timing: PTP (IEEE 1588) design and troubleshooting, grandmaster placement, boundary and transparent clock concepts, sync quality monitoring against microsecond budgets.
Fabrics for AI/HPC: PFC and lossless Ethernet tuning, link training, SerDes and optics basics (current generation high speed optics), awareness of how skew and tail latency hit collective communication.
Optical and physical plant: fiber types and connectors, optical circuit switching concepts, commercial traffic generators, cable validation, fiber characterization from hardware timestamps.
Access security: port ACLs, 802.1X/MAB flows, RADIUS, quarantine and remediation VLANs, device onboarding and registration.
Multi-vendor fluency: deep on at least one datacenter NOS (Arista, Cisco, Juniper, SONiC) with demonstrated ability to pick up others, including API driven operation with rollback.
Troubleshooting under pressure: packet level debugging, hardware timestamp analysis, method of procedure authorship, maintenance windows, rollback planning.
Required: automation and software
Python for network automation: device APIs, text templating, input validation, operator CLI tooling.
Templating: Jinja2 template design (reuse across sites and roles, safe defaults, per-port rendering from inventory).
Config as code: distributed revision control fluency (branch, review, land, revert), config review culture, CI gates on generated config (diff review, automated land criteria).
Intent modeling: declarative source of truth thinking where inventory plus intent renders device config, and generated config is never hand edited.
Linux networking: interfaces, routing tables, namespaces, capture tooling, service management, cookbook style config management (chef).
Data hygiene: inventory accuracy (roles, port maps, serial to port bindings), drift detection, reconciliation loops.
Required: operations and delivery
Lab or datacenter operations: oncall rotations, incident response, break/fix prioritization, vendor RMAs, multi-site coordination.
Change discipline: method of procedure authorship and peer review, blast radius control, pilot then scale.
Cost awareness: circuit inventories, carrier coordination for turn-down, tracking recurring spend to zero with chargeback confirmation.
Cross-functional delivery: shared ownership with platform teams (provisioning, observability), hardware bring-up teams, and lab customers, including consultative support outside the plan of record.
Preferred
AI/HPC cluster networking: scale-up versus scale-out fabric design, RDMA capable fabrics, collective sensitivity to timing.
Large fiber plant or optical switching operations.
IPv6-first or IPv6-only environments at scale.
Observability: time series metrics, topology visualization, low-noise alert design.
Vendor evaluation: switch selection criteria, optics compatibility testing.
What predicts success here
This team works through internal platforms that all have direct industry analogues. We hire for the transferable skill; ramp covers the internal name. Git-backed hierarchical config management plus Jinja (Ansible/Salt/Puppet/Chef experience transfers), policy driven intent automation (declarative config pipelines transfer), graph modeled network state with diff-before-push (topology as data transfers), NetBox class inventory and asset sources of truth, in-house orchestration and topology visualization (any NMS you have operated or extended transfers), and Git class revision control with daily diff review.