← Toutes les offres

Swan fournit aux entreprises une infrastructure de services financiers intégrés, leur permettant de proposer des comptes, des cartes et des paiements au sein de leurs propres produits.

Détails de l’offre

La description complète publiée par Swan.

Voir l’offre originale ↗

Description de l’offre

Working within Swan’s Core Infrastructure team, you will help ensure the reliability, scalability, security, and performance of the platforms that support our financial services. You will take ownership of well-scoped services and operational incidents, improve observability, automate repetitive tasks, and collaborate closely with development, product, and security teams. This is an independent engineering role for someone who has developed solid operational foundations and is ready to take greater ownership of production systems.

You will contribute to incident response, infrastructure improvements, service design reviews, and the continuous improvement of our reliability practices.

Main responsibilities

On a daily basis, you will: Act as a primary responder for well-understood production incidents and participate independently in the on-call rotation. Assess the impact of incidents, including transaction volume affected, potential revenue impact, and implications for data integrity. Investigate operational issues using logs, metrics, dashboards, and distributed tracing, then document clear incident updates for stakeholders. Contribute to postmortems, update runbooks, identify recurring incident patterns, and suggest preventive measures.

  • Create and maintain dashboards, alerts, and basic service-level indicators for the services you support. Tune alert thresholds to reduce noise and improve the quality of operational signals. Participate in system design reviews, with a particular focus on reliability, operability, failure modes, and production readiness.
  • Implement reliability improvements such as health checks, retries with exponential backoff, circuit breakers, and appropriate monitoring.
  • Manage cloud resources and contribute to Infrastructure as Code using tools such as Terraform. Write automation scripts and small internal tools in Bash, Python, or Go to reduce manual toil and improve operational efficiency. Contribute to CI/CD pipelines and automate routine maintenance tasks such as backup verification, certificate renewal, and log management. Participate in infrastructure code reviews and help maintain high standards for safe, repeatable changes.
  • Support security and compliance activities, including PCI DSS controls, ISO 27001 initiatives, security remediation, data classification, and encryption requirements.
  • Monitor resource utilisation, provide basic capacity forecasts, and implement practical cost optimisation measures such as rightsizing resources and removing unused infrastructure.
  • Collaborate with development, product, and security teams to improve the resilience and operability of services.
  • Provide clear handovers, maintain high-quality documentation, and communicate technical topics effectively to both technical and non-technical stakeholders. Use approved AI tools responsibly to support tasks such as code generation, documentation, and log analysis, while validating outputs and protecting sensitive information. Your team Core Infrastructure is responsible for building and operating the foundations that enable Swan’s products to remain reliable as the business grows. We work closely with development and other technical teams to improve system resilience, operational efficiency, and customer experience. We value ownership, pragmatism, knowledge sharing, and open communication. Engineers are encouraged to challenge ideas constructively, document what they learn, and continuously improve the way we build and operate services. You will work in a supportive environment where reliability is a shared responsibility and where operational excellence is developed through collaboration. Together alongside Engineering Productivity, our squad constitutes the broader Platform Engineering team.

✨ You’re a great match if: You have typically 2 to 4 years of experience in Site Reliability Engineering, DevOps, platform engineering, infrastructure engineering, software engineering, or a related field. You have hands-on experience supporting production services and participating in an on-call rotation. You can independently respond to well-understood incidents, follow escalation procedures, and contribute to postmortems and runbook improvements. You are comfortable working with logs, metrics, dashboards, alerting, and basic distributed tracing.

You understand the Four Golden Signals: latency, traffic, errors, and saturation. You have experience creating dashboards and meaningful alerts, and understand the fundamentals of SLIs and SLOs. You have practical experience with cloud infrastructure, ideally AWS, including compute, storage, networking, and managed services. You have experience with Infrastructure as Code, particularly Terraform, CloudFormation, or equivalent tools. You can write automation scripts in one or more languages such as Bash, Python, or Go.

You understand CI/CD practices and have contributed to deployment or infrastructure automation. You have a working understanding of high availability, fault tolerance, redundancy, health checks, retries, circuit breakers, and failure recovery. You are familiar with event-driven architectures and message queuing technologies such as Kafka or AWS SQS. You understand the operational implications of distributed systems, including consistency, availability, and partition tolerance.

You have an awareness of security and compliance requirements relevant to financial services, such as PCI DSS, ISO 27001, encryption, least privilege, and data classification. You are comfortable monitoring resource usage, thinking about capacity, and identifying basic cloud cost optimisation opportunities. You communicate clearly, document your work thoroughly, and collaborate effectively across technical and non-technical teams.

Nice to have

AWS certification, such as AWS Certified Solutions Architect, Associate level. Experience with Kubernetes and container orchestration. Experience with observability platforms such as Grafana, Datadog, or similar. Experience operating services that process financial transactions or other highly sensitive data. Experience improving service-level objectives, performance, or capacity for production systems. Familiarity with Go or another programming language used for internal tooling. Our ideal teammate: Empathetic. Skilled. Frank.

We love to challenge each other, and we leave our egos at the door. It’s okay if you don’t tick all the boxes - don’t let imposter syndrome prevent you from applying! 🙌 Swan is committed to providing a caring work environment for all employees, regardless of age, sex, disability, sexual orientation, race, religion, or belief. When it comes to recruitment, we’re interested in your work experience, skills, and overall personality.

Because diversity makes the workplace stronger and is necessary for Swan’s success, we are intensifying efforts to incorporate concrete actions to help us improve in this area.

A 30-min call with our Talent Acquisition Manager, to get to know you, understand your career expectations and answer your questions An interview with our Platform Director A tech test & peer interview An interview with our CTO

Prérequis

  • expérience pratique du support de services en production
  • participation à une astreinte
  • création de tableaux de bord et d’alertes
  • compréhension des SLI et SLO
  • expérience de l’infrastructure cloud
  • expérience de l’Infrastructure as Code
  • compréhension des pratiques CI/CD
  • communication claire et documentation rigoureuse