Site Reliability Engineer (SRE) – AWS, Terraform & CI/CD
Location: Remote – LATAM
Employment Type: Full-time, Contract
Compensation: Up to $8,300/month
FitNext Exclusive: Yes
Company Overview
Join a growing technology company building a scalable, secure, and reliable platform. The engineering team is focused on infrastructure consolidation, deployment automation, and operational excellence, creating the foundation for continued growth.
The Challenge
We are looking for a hands-on Site Reliability Engineer to build and maintain the infrastructure, automation, and reliability practices that support a production-grade technology platform.
You will play a key role in eliminating manual deployment processes, improving system observability, and building reliable cloud infrastructure. You will also contribute to infrastructure automation, disaster recovery, data platform improvements, and SOC 2 Type 2 compliance.
This role is ideal for an engineer who enjoys solving complex infrastructure challenges, taking ownership, and building scalable systems in a fast-moving environment.
Must-Have
- Strong hands-on experience with Terraform or equivalent Infrastructure as Code (IaC) tools, managing infrastructure across multiple environments.
- Solid experience with AWS, including services such as EC2, RDS, S3, IAM, VPC, Lambda, and ECS/EKS.
- Proven experience building and maintaining CI/CD pipelines, preferably with GitHub Actions or similar tools.
- Strong Python scripting and automation skills, along with Bash.
- Deep knowledge of Linux administration, networking fundamentals, and cloud security best practices.
- Experience with monitoring, logging, and alerting tools such as Datadog, Prometheus/Grafana, ELK, or CloudWatch.
- Understanding of SRE principles, including system reliability, incident response, disaster recovery, and blameless post-mortems.
- Strong English communication skills and the ability to collaborate with technical and non-technical stakeholders.
- Comfortable working autonomously, troubleshooting complex issues, and making pragmatic decisions in ambiguous environments.
- Willingness to participate in an on-call rotation after an appropriate training period and support occasional scheduled after-hours maintenance.
Nice to Have
- Experience with Snowflake, including database administration, permissions, and environment isolation.
- Familiarity with dbt, SQL, data pipelines, and data quality testing frameworks.
- Experience with GitOps and automated provisioning of development, staging, production, sandbox, and UAT environments.
- Knowledge of SLIs, SLOs, error budgets, and chaos engineering.
- Experience implementing security controls for SOC 2 Type 2 or similar compliance frameworks.
- Familiarity with Clojure or other programming languages beyond Python.
- Experience designing disaster recovery strategies and conducting recovery drills.
What You'll Do
- Design and maintain Terraform configurations for AWS and Snowflake infrastructure.
- Build automated infrastructure provisioning workflows using GitOps.
- Develop CI/CD pipelines that minimize manual intervention and enable reliable production deployments.
- Improve observability through monitoring, logging, alerting, dashboards, and service reliability metrics.
- Design disaster recovery plans and participate in regular recovery exercises.
- Collaborate with Data Engineering to migrate Snowflake pipelines to dbt and implement data quality checks.
- Lead infrastructure incident response, conduct blameless post-mortems, and drive continuous reliability improvements.
- Support security initiatives, IAM hardening, secrets management, and SOC 2 Type 2 compliance.
- Maintain clear infrastructure documentation, runbooks, and operational procedures.
Hiring Process
Details will be shared with shortlisted candidates.
Interested? Apply now to be considered for this opportunity!