Woodfrog is a Data & AI Engineering company based in Pune, India, focused on helping organizations transform data into reliable insights and production-ready AI solutions. We build modern data platforms, AI-powered applications, analytics solutions, and cloud infrastructure that enable businesses to make faster, data-driven decisions. Founded in 2023, Woodfrog specializes in data engineering, analytics, AI agents, governance, and enterprise AI deployments while delivering scalable, secure, and high-performance solutions for clients across industries.
About the Opportunity
This role is with a Woodfrog client that is building next-generation AI/ML platforms and cloud-native infrastructure to support enterprise-scale applications. The client focuses on delivering secure, scalable, and automated cloud solutions while enabling modern AI workloads through robust DevOps and MLOps practices.
As a Lead DevOps Engineer (AI/ML Platform), you'll play a key role in designing and managing AWS infrastructure, leading Kubernetes-based deployments, building CI/CD automation, and collaborating with Engineering and AI/ML teams to deliver reliable, production-ready platforms. If you're passionate about cloud infrastructure, automation, and building scalable systems that power cutting-edge AI applications, we'd love to hear from you.
About the Role
You'll play a key role in building and scaling the cloud infrastructure that powers enterprise AI/ML platforms. Expect hands-on work across cloud architecture, infrastructure automation, Kubernetes orchestration, and CI/CD engineering. You'll collaborate closely with engineering and AI/ML teams to deliver secure, reliable, and production-ready deployments while driving DevOps best practices across the platform.
Key Responsibilities
Design, build, and manage scalable AWS cloud infrastructure for enterprise AI/ML platforms.
Develop and maintain Infrastructure as Code (IaC) using Terraform to automate cloud provisioning and deployments.
Build, manage, and optimize Kubernetes (EKS) clusters, Helm deployments, and containerized applications.
Design and enhance Jenkins CI/CD pipelines to automate build, test, release, and deployment workflows.
Monitor, troubleshoot, and optimize production infrastructure, ensuring high availability, security, and reliability.
Collaborate with Engineering, QA, and AI/ML teams to streamline deployments and improve platform performance.
Implement monitoring, logging, and observability solutions to proactively identify and resolve system issues.
Maintain Linux-based environments, automate operational tasks using Bash/Python scripting, and follow DevOps best practices.
Drive infrastructure scalability, security, and operational excellence while mentoring team members and promoting engineering best practices.Requirements
Bachelor's or Master's degree in Computer Science, Information Technology, or a related field.
8+ years of hands-on experience in DevOps, Cloud Engineering, or Platform Engineering.
Strong expertise in AWS services, including VPC, EKS, IAM, RDS, and cloud networking.
Proven experience with Infrastructure as Code (Terraform) and cloud automation.
Hands-on experience with Kubernetes (EKS), Docker, Helm, and container orchestration.
Strong knowledge of Jenkins, CI/CD Pipeline-as-Code, and deployment automation.
Proficiency in Linux administration and scripting using Bash and/or Python.
Experience with Git workflows, JFrog Artifactory, release management, and production support.
Working knowledge of PostgreSQL and cloud-based production environments.
Strong understanding of monitoring, logging, observability, and infrastructure reliability.
Excellent problem-solving, troubleshooting, and communication skills.
Comfortable working in a fast-paced startup environment with a high level of ownership and collaboration.Nice To Have
Experience with AWS Bedrock, Kubeflow, MLflow, Langfuse, LiteLLM, Weaviate, or NebulaGraph.
Exposure to GenAI, RAG applications, vector databases, and AI/ML platform deployments.
Experience with GitOps, ArgoCD, HashiCorp Vault, and Kubernetes-based AI workload deployments.Why Join Us
Competitive compensation package with the opportunity to work on high-impact enterprise AI/ML platforms.
Lead the design and automation of cloud infrastructure powering next-generation AI applications.
Work with modern technologies including AWS, Kubernetes, Terraform, Jenkins, and MLOps tools.
Collaborate with experienced engineers and AI/ML teams to build scalable, production-ready platforms.
Take ownership of critical infrastructure decisions in a fast-paced, innovation-driven environment.
Grow your career by working on challenging cloud, DevOps, and AI infrastructure projects with significant learning and leadership opportunities.