Infrastructure Engineer / SRE

Infrastructure Linux · Networking Moscow · On-site Full-time

The role

You will own the physical and logical infrastructure that keeps our trading systems running with zero unplanned downtime during market hours. This includes co-location server management, network topology and cross-connect provisioning, Linux kernel tuning for latency, and the observability stack that enables the engineering team to diagnose issues at nanosecond resolution.

This is a hands-on role — you will rack and cable servers, configure switches, write Ansible playbooks, and be on-call when something breaks. There is no separation between "infra" and "ops" here.

Responsibilities

  • Manage server hardware at co-location facilities: provisioning, hardware fault isolation, component replacement coordination.
  • Own network topology and cross-connect infrastructure; maintain routing tables, VLAN config, and BGP peering with exchange feeds.
  • Tune Linux systems for low-latency operation: CPU isolation, IRQ affinity, NUMA binding, huge pages, transparent huge page disabling.
  • Maintain PTP clock synchronisation infrastructure across all co-located and office systems.
  • Build and maintain the observability platform: metrics collection, alerting, latency dashboards, and post-trade analysis pipelines.
  • Automate deployment and configuration management using Ansible and custom tooling; maintain infrastructure-as-code discipline.
  • Participate in incident response and post-mortem process; write detailed root-cause analyses.

Requirements

  • 5+ years of Linux systems administration or SRE experience in a performance-critical environment.
  • Deep knowledge of Linux networking: tc/qdisc, ethtool, kernel networking stack, NIC driver tuning.
  • Experience with data-centre networking: switch configuration, VLAN design, LACP, BGP basics.
  • Proficiency with infrastructure automation: Ansible, Terraform, or equivalent.
  • Solid scripting ability in Python and Bash; comfort reading C or Go is a plus.
  • Experience with time-series observability stacks (Prometheus, InfluxDB, Grafana, or custom solutions).
  • Ability to read kernel traces and system-level profiling output (perf, ftrace, eBPF).

Nice to have

  • Prior experience at a co-location facility, exchange, or high-frequency trading firm.
  • Familiarity with DPDK or RDMA network stacks from an operational perspective.
  • Experience managing PTP/IEEE-1588 clock synchronisation in production.
  • Knowledge of FPGA PCIe device management and driver operations.

How to apply

Send a CV and a brief description of the most complex infrastructure problem you have solved to:

research@klarnet-trading.ru

Subject line: SRE — [Your name]

Specifics matter more than job titles — tell us about the incident you diagnosed at 3 am and what you learned.

Contact →