Start Your Search Here

Job Search

Luxoft

Puerto Vallarta / Global

Linux Systems Engineer

  • $300.000 - $450.000

Job Description

Project description

We are looking for a Linux Systems Engineer for Large US hedge fund to keep critical Linux-based research and platform infrastructure reliable, secure, and efficient. This role supports production systems across Linux servers, Kubernetes platforms, scheduler-backed compute environments, and user access services. The engineer will troubleshoot incidents, automate operational work, improve observability, and partner with cross-functional teams to deliver stable infrastructure changes.

Responsibilities

- Operate and support Linux servers and shared infrastructure used by research and platform teams.

- Troubleshoot production issues involving system performance, availability, access, configuration, and networking.

- Support Kubernetes and other container-based platforms, including node health, service behavior, and rollout activities.

- Support scheduler-backed compute environments such as Slurm, including node readiness, maintenance, and incident recovery.

- Improve user-facing Linux access services such as SSH, shared shell environments, and session-based platforms.

- Manage OS lifecycle work: provisioning, patching, kernel and package updates, and hardening.

- Build scripts and automation in Python, Bash, or similar tools to reduce manual work and improve reliability.

- Use configuration management and version-controlled workflows to implement infrastructure changes safely.

- Enhance monitoring, alerting, documentation, and operational processes.

- Participate in incident response and occasional on-call support.

SKILLS

Must have

- 3+ years of experience in Linux systems engineering, SRE, DevOps, or infrastructure support.

- Strong Linux administration skills, including systemd, package management, permissions, filesystems, loganalysis, and performance troubleshooting.

- Good understanding of networking fundamentals such as DNS, NTP/PTP, routing, and general host connectivity.

- Experience with automation, scripting, and operational tooling.

- Familiarity with Kubernetes, virtualization, or clustered platforms.

- Experience with configuration management or infrastructure-as-code tools such as Ansible, Salt, or Terraform.

- Ability to troubleshoot production issues methodically and communicate clearly during incidents.

- Experience with Git-based workflows and maintainable documentation.

- Hands-on, practical problem solver with a strong ownership mindset.

- Comfortable working close to production and balancing support with continuous improvement.

- Collaborative communicator who works well across compute, storage, networking, and application teams.

Nice to have

- Experience with Slurm, HPC-style environments, GPU infrastructure, or researcher-facing Linux platforms.

- Working familiarity with shared storage clients such as NFS, autofs, or GPFS / IBM Storage Scale from a host and application perspective.

- Experience with observability tools such as Prometheus, Grafana, or equivalent platforms.

- Exposure to identity and access services such as LDAP, Kerberos, SSSD, or PAM.

- Exposure to on-premises datacenter operations, hardware lifecycle support, or vendor escalations.

- Interest in using AI/ML techniques for infrastructure optimization, anomaly detection, or predictive operations.

#J-18808-Ljbffr

Aplicar Now

Similar Opportunities

View all jobs