Full job description
Join Vibe's Infra Platform team responsible for hosting, provisioning, observability, CI/CD, internal tooling, security, and compliance. Report to the Lead Platform Engineer and help scale infrastructure supporting 600k QPS, latency under 10ms, and over 5 PB storage. Responsibilities include maintaining 99.99% uptime, proactive infrastructure maintenance, incident response, designing scalable infrastructure components, managing infrastructure as code (Terraform), improving developer experience with tooling in Go or Python, supporting ML infrastructure and real-time streaming systems, and enforcing security best practices including SOC2 compliance. Requirements include hands-on experience with large-scale production infrastructure, Terraform expertise, fluency in Go or Python, deep CI/CD and observability knowledge, incident response leadership, and effective use of AI in engineering. Benefits include variable pay, hybrid flexibility (Paris office), full health insurance, meal vouchers, annual offsite, and quarterly tech syncs.
What you'll do
- Own reliability across hosting, provisioning, network, and compute, targeting 99.99% uptime
- Proactively maintain infrastructure, track drift, and lead migrations before they become incidents
- Respond to production issues quickly, communicate clearly, and fix root causes for lasting improvements
- Design and implement core infrastructure components that scale significantly
- Make build-vs-buy decisions and evolve systems as business needs shift
- Manage infrastructure as code across different providers
- Improve developer experience by building shared tooling such as internal CLIs, shared libraries, and automation scripts in Go or Python
- Partner with development teams to help them ship faster without creating technical debt
- Support ML infrastructure for frequent model retraining on large datasets
- Improve compute efficiency and model serving performance to meet inference latency targets
- Build, scale, and operate the real-time streaming platform
- Enforce best practices around permissions, secret handling, and network security
- Support SOC2 compliance work and embed security into new projects from day one
- Balance risk trade-offs pragmatically to ensure protection without paralysis
Requirements
- Hands-on experience operating production infrastructure at meaningful scale, with strong instincts around reliability, resilience, and performance under load
- Strong experience with infrastructure as code, especially Terraform, with ownership of the full lifecycle from implementation to continuous improvement
- Fluency in at least one systems-oriented language, ideally Go or Python, with the ability to build automation and operational tooling
- Deep experience with CI/CD, observability, and production operations, including metrics, logs, traces, alerting, and debugging live systems
- Comfortable leading incident response, improving service reliability, and driving root cause resolution
- Able to support product and engineering teams as a trusted infrastructure partner
- Uses AI effectively in day-to-day engineering, with strong judgment about when it adds leverage and when deeper manual work and critical thinking matter more
Tech stack
TerraformGoPythonCI/CDobservabilityinfrastructure as codereal-time streamingML infrastructureautomation scripts
Benefits
Variable pay based on objectivesHybrid flexibility (office in the Heart of Paris)Full health insurance coverage via AlanMeal vouchers via SwileAnnual offsite for the whole teamQuarterly Tech Syncs with Engineering and Product teams worldwide