About this role
Phone numbers and emails in this ad are masked until you log in.
auto_translated_note
We are looking for an MLOps Engineer with a strong interest in automation, cloud computing and efficient deployment of intelligent platforms. If you are motivated by improving productive environments, guaranteeing scalable and reliable systems, and working with cutting-edge technologies in Machine Learning, AI Agents and Observability, this opportunity is for you.🚀 What will your role be? As an MLOps Engineer, you will be responsible for connecting the world of Data and ML with the infrastructure and operation in production, ensuring that models and agents function robustly and efficient.Your main responsibilities will include: Design, implement and operate scalable cloud infrastructures in AWS, GCP or Azure, aimed at data and ML workloads.
Automate infrastructures and deployments using Terraform, Ansible, Helm and Kubernetes. Create and maintain CI/CD pipelines for the continuous deployment of models, AI agents and observability platforms. Manage containers and orchestration with Docker and Kubernetes (EKS), integrating data, AI and backend services.
Build and maintain reproducible training, validation and inference pipelines (Airflow, MLflow, DVC, Spark). Implement and optimize monitoring, logging and observability solutions (Prometheus, Grafana, ELK/EFK, OpenTelemetry). Ensure the security, availability and resilience of cloud platforms.
Integrate the complete Data → ML → Deployment → Monitoring cycle, working closely with Data, ML, DevOps and Backend teams. Collaborate with Machine Learning teams to package and bring AI models and agents to production. (Advanced Plus) Participate in the design of hybrid pipelines that combine traditional ML, generative AI and LLM-based agents (LangChain, LangGraph, CrewAI). 🎯
Requirements
We are looking for someone with: Training in Computer Engineering, Software, Telecommunications or related disciplines. More than 3 years of experience in automation, deployment or management of cloud infrastructure. Strong experience in infrastructure as code and automation (Terraform, Helm, Ansible).
Advanced knowledge of Linux, networks and distributed systems. Hands-on experience in CI/CD pipelines (GitLab CI, Jenkins, ArgoCD, FluxCD). Docker and Kubernetes proficiency.
Experience with model and data management tools such as MLflow, DVC or Vertex AI. Knowledge of monitoring, logging and observability. Experience with SQL and NoSQL databases (PostgreSQL, TimescaleDB, MongoDB).
Ability to diagnose and resolve incidents in high-traffic productive environments, optimizing costs and performance. Good communication and teamwork skills in multidisciplinary environments. 🌟 We especially value if you also have experience with messaging and streaming systems such as Kafka, Redpanda or Benthos. Knowledge of serverless and event-driven architectures (AWS Lambda, SNS/SQS).
Experience in advanced observability and model performance metrics. Knowledge of SRE (Site Reliability Engineering) practices. Experience with modern MLOps platforms (Kubeflow, MLflow).
Familiarity with automated deployments and good DevOps practices. Knowledge of infrastructure for generative AI (GPU, optimized containers, RAG serving). Advanced technical English, written and spoken, to collaborate with international teams and partners.
Originally posted on Himalayas
Community Q&A
Anyone worked here? Ask before you apply.
No threads yet for this job or company.