Arlequin AI développe HuDex, une plateforme d’analyse de données multilingues et non structurées, et mène des travaux de recherche en apprentissage profond topologique. Ses outils s’adressent aux entreprises et aux organismes publics.
Site Reliability Engineer
Am I a fit — voir ma compatibilitéArlequin AI recherche un Senior Site Reliability Engineer pour renforcer son équipe Platform à Paris et opérer une plateforme cloud-native souveraine hébergée sur Scaleway. Le poste couvre la fiabilité de la production, l’infrastructure as code, Kubernetes, le GitOps et l’observabilité, avec une pratique de Bash et Python. Le travail peut s’effectuer en mode hybride, totalement à distance ou sur site, selon une organisation flexible.
Repères sur Arlequin AI
- Secteur
- Logiciels et Internet
- Modèle
- B2B
- Siège
- Paris
- Domaine officiel
- arlq.ai
- Dernière levée
- Série A · 2026-09-01
- Montant annoncé
- 31 000 000 $
- Offres ouvertes
- 10
Investisseurs mentionnés
Détails de l’offre
La description complète publiée par Arlequin AI.
Description de l’offre
← Back to roles
PLATFORM
Site Reliability Engineer Ensure the availability, performance, and scalability of Arlequin's sovereign, cloud-native platform
5+ YEARS, SENIOR
FULL-TIME
PARIS (HYBRID / FULL-REMOTE / ON-SITE FLEXIBLE)
About Arlequin
Arlequin AI is both a topological deep learning research lab and an AI platform. Our first product, HuDex, converts massive volumes of raw, unstructured, multilingual data into strategic decisions in minutes instead of days, providing an operational advantage for government agencies and businesses.
The role
We have chosen a cloud-native, sovereign, end-to-end automated infrastructure, hosted on Scaleway. Joining us means taking part in building and operating a reliable, secure, and observable platform serving our research, product, and development teams.
The Platform team owns the technical foundations Arlequin AI runs on: DevOps, DevSecOps, MLOps, DataOps, FinOps, security and compliance, compute, and IT. Our role is not to build infrastructure for its own sake — we build self-service tools so engineers can ship without waiting on us, a Forward Deployed Engineer can deploy a client without our help, C-levels understand the real cost of what they sell, and researchers don’t have to deal with engineering questions to run their experiments. Today we are a team of 3, and we are looking for people to help structure the team.
As a Senior SRE, you join the Platform team to ensure the availability, performance, and scalability of our platforms, while giving development teams the tooling to be autonomous. You work on both the run and the build side, with a genuine culture of toil reduction.
Your responsibilities
Reliability & production: define, instrument, and track SLIs / SLOs / SLAs; continuously improve resilience (capacity planning, load testing, chaos engineering, disaster recovery, eliminating SPOFs); reduce toil through automation and self-service; handle incidents and run blameless post-mortems; support the scaling of training, inference, and scientific computing workloads
Infrastructure & IaC: design, deploy, and maintain infrastructure on Scaleway; industrialize IaC with OpenTofu and Terragrunt; administer and evolve production Kubernetes clusters
GitOps & delivery: operate and evolve continuous deployment following GitOps principles with ArgoCD; standardize CI/CD pipelines and release workflows; manage configuration and secrets securely
Observability: improve and maintain our LGTM stack (Loki, Grafana, Tempo, Mimir); design dashboards and a relevant alerting policy with Alertmanager; promote OpenTelemetry observability and train engineering teams on it
Culture & collaboration: spread SRE / DevOps best practices across teams; document the infrastructure and runbooks; contribute to architecture decisions and the platform’s technical roadmap
On-call: no rotation today, incidents are handled during business hours. When a rotation becomes necessary, it will never exceed one week on-call out of five, it will be paid, and you’ll take part in designing it
Stack
Cloud & orchestration: Scaleway, Kubernetes
IaC & GitOps: OpenTofu, Terragrunt, ArgoCD
Networking: Cilium
Observability: Loki, Grafana, Tempo, Mimir (LGTM), Prometheus, Alertmanager
Languages: Bash, Python
What we’re looking for
5+ years of experience in SRE / DevOps / Infrastructure, including significant experience with mission-critical production
Advanced command of Kubernetes in production environments
Solid experience with IaC (Terraform / OpenTofu, ideally Terragrunt)
Hands-on GitOps practice (ArgoCD or Flux)
Strong observability culture (ideally the LGTM stack)
Comfortable with scripting / automation (Bash, Python) and solid Linux fundamentals
Autonomy and a strong sense of ownership over the production scope, excellent communication, a feedback culture, technical curiosity, and pragmatism
Bonus
Hands-on experience with Scaleway
Knowledge of Cilium & service mesh
FinOps awareness
Process
First-fit interview (30 min)
Technical interview with the Platform team (coding + design)
Meeting with the hiring manager
Offer
Package
Competitive compensation including equity stake
Flexible remote policy: hybrid / full-remote / full on-site
Training, conference, and certification budget, with time dedicated to CNCF/LF open source contributions
Prérequis
- Advanced command of Kubernetes in production environments
- Solid experience with IaC
- Hands-on GitOps practice
- Strong observability culture
- Comfortable with scripting / automation
- solid Linux fundamentals
- Autonomy
- excellent communication
- technical curiosity
- pragmatism
Avantages mentionnés
- equity stake
- Flexible remote policy
- Training, conference, and certification budget
- time dedicated to CNCF/LF open source contributions