Jobs / Ora***

Senior Core Infrastructure Engineer

Ora*** · Nashville, TN, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
Nashville, TN, United StatesOnsite
Remuneration
Not specified
Location
Nashville, TN, United States
Visa sponsorship
Sponsors visa

Job description

Designs, implements, and optimizes components in distributed systems with an emphasis on scalability, resiliency, and operability. Delivers features and load/performance tests; leverages data plane platforms and distributed state tools for high-volume retrieval, storage, and processing; and reviews peers’ implementations for scalability compliance. Builds fault-tolerant paths (redundancy, replication, automatic failover), applies recovery‑oriented principles, and implements retries, circuit breakers, and timeouts. Proactively detects and mitigates issues via tests, alarms, dashboards, and telemetry; authors runbooks and participates in incident response and RCAs. Implements standard replication and synchronization, develops automation/IaC for troubleshooting and maintenance, and applies advanced security controls (encryption, access, remediation) while ensuring change, compliance, and documentation standards are met. Senior Core Infrastructure Engineer, DNS Data Plane We are seeking a Senior Core Infrastructure Engineer to join our DNS Data Plane team. You will design, build, and operate high-performance, highly reliable DNS services that power critical infrastructure at a global scale. This role focuses on performance-sensitive systems, distributed networking, and operational excellence, balancing correctness, latency, and resiliency. Key Responsibilities Design and implement DNS data plane components across authoritative and recursive request paths, emphasizing low-latency, high-throughput processing. Build services and libraries in Go, C, Java, or Python, selecting the appropriate language and approach for each component’s performance and operational requirements. Optimize CPU, memory, and network I/O through efficient event loops, concurrency models, socket tuning, profiling, and benchmarking. Own data plane components end to end, including architecture, implementation, testing, safe rollout, observability, capacity planning, and ongoing operational health. Develop and maintain metrics, logs, traces, SLOs, dashboards, alerts, and practical on-call playbooks. Improve reliability through fault-tolerant design, graceful degradation, safe deployment practices, and proactive capacity planning. Collaborate with control plane, security, SRE, and operations teams to deliver end-to-end DNS platform improvements. Provide technical leadership through architecture reviews, code reviews, and mentorship, raising engineering standards across the team. Take ownership of production outcomes by proactively identifying and driving performance, reliability, security, and operational improvements. Required Qualifications At least five years of experience building, managing, or operating large-scale distributed infrastructure in SaaS, cloud, or hybrid-cloud environments. At least three years of hands-on experience with Linux operating systems. Strong experience building and operating production-grade distributed systems or network services. Proficiency in one or more of Go, C, Java, or Python, with the ability and interest to work across languages when needed. Solid understanding of Linux systems, networking fundamentals, TCP/UDP, sockets, concurrency, and performance profiling. Demonstrated experience designing highly available systems and delivering safe production rollouts using practices such as canaries, progressive deployments, and feature flags. Strong ownership and sound technical judgment, with the ability to navigate ambiguity and drive complex problems to durable resolution. Desired Skills Deep DNS knowledge, including relevant RFCs, DNS over UDP and TCP, EDNS(0), DNSSEC, caching behavior, zone transfers, rate limiting, and load balancing. Experience with high-performance systems patterns such as epoll/kqueue, asynchronous I/O, low-lock or lock-free data structures, and packet processing. Familiarity with DDoS mitigation, abuse prevention, and traffic engineering for L4 and L7 services. Experience with containers and orchestration platforms such as Kubernetes, including service meshes and network policies where relevant. Infrastructure-as-code experience, particularly with Terraform. Strong testing discipline, including unit, integration, fuzz, property-based, performance, and chaos testing. Experience operating mission-critical, 24/7 services. Experience using AI-assisted development tools to accelerate coding, testing, debugging, and documentation while maintaining strong engineering judgment, ownership, code quality, and security standards. What Success Looks Like You deliver measurable improvements in latency, throughput, reliability, and operational efficiency. You take end-to-end ownership of the systems you build, from initial design through long-term production health. You evolve the DNS platform through clean architecture, strong operational readiness, and secure-by-design practices. You identify risks and opportunities early and drive pragmatic improvements beyond the immediate task. You raise team standards through mentorship, thoughtful technical leadership, and a culture of operational excellence.

Skills

GoJavaTerraformLinuxKubernetesPython