Infrastructure Engineer / SRE
Infrastructure
Linux · Networking
Moscow · On-site
Full-time
The role
You will own the physical and logical infrastructure that keeps our trading systems running with zero unplanned downtime during market hours. This includes co-location server management, network topology and cross-connect provisioning, Linux kernel tuning for latency, and the observability stack that enables the engineering team to diagnose issues at nanosecond resolution.
This is a hands-on role — you will rack and cable servers, configure switches, write Ansible playbooks, and be on-call when something breaks. There is no separation between "infra" and "ops" here.
Responsibilities
- Manage server hardware at co-location facilities: provisioning, hardware fault isolation, component replacement coordination.
- Own network topology and cross-connect infrastructure; maintain routing tables, VLAN config, and BGP peering with exchange feeds.
- Tune Linux systems for low-latency operation: CPU isolation, IRQ affinity, NUMA binding, huge pages, transparent huge page disabling.
- Maintain PTP clock synchronisation infrastructure across all co-located and office systems.
- Build and maintain the observability platform: metrics collection, alerting, latency dashboards, and post-trade analysis pipelines.
- Automate deployment and configuration management using Ansible and custom tooling; maintain infrastructure-as-code discipline.
- Participate in incident response and post-mortem process; write detailed root-cause analyses.
Requirements
- 5+ years of Linux systems administration or SRE experience in a performance-critical environment.
- Deep knowledge of Linux networking: tc/qdisc, ethtool, kernel networking stack, NIC driver tuning.
- Experience with data-centre networking: switch configuration, VLAN design, LACP, BGP basics.
- Proficiency with infrastructure automation: Ansible, Terraform, or equivalent.
- Solid scripting ability in Python and Bash; comfort reading C or Go is a plus.
- Experience with time-series observability stacks (Prometheus, InfluxDB, Grafana, or custom solutions).
- Ability to read kernel traces and system-level profiling output (perf, ftrace, eBPF).
Nice to have
- Prior experience at a co-location facility, exchange, or high-frequency trading firm.
- Familiarity with DPDK or RDMA network stacks from an operational perspective.
- Experience managing PTP/IEEE-1588 clock synchronisation in production.
- Knowledge of FPGA PCIe device management and driver operations.