Senior Platform Engineer
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
By continuing, you agree to our Terms & Privacy Policy.
About The Team
Our Platform Engineering team builds and operates the core backend systems that power CoCounsel's AI agent platform — the infrastructure that lets legal AI agents run reliably at scale. We own the services that sit between product-facing chat/agent experiences and the AI runtime layer, the CI/CD and deployment infrastructure that ships them safely, and the internal developer tooling that other engineering teams build on. We work closely with product, applied-AI, and platform teams to ship AI-driven capabilities faster and more reliably across the business.
About The Team
Our Platform Engineering team builds and operates the core backend systems that power CoCounsel's AI agent platform — the infrastructure that lets legal AI agents run reliably at scale. We own the services that sit between product-facing chat/agent experiences and the AI runtime layer, the CI/CD and deployment infrastructure that ships them safely, and the internal developer tooling that other engineering teams build on. We work closely with product, applied-AI, and platform teams to ship AI-driven capabilities faster and more reliably across the business.
As a Senior Platform Engineer, you will design and deliver backend services with growing scope and autonomy, contribute to architectural decisions, and help raise engineering standards across the team.
About The Role
- Design and deliver backend and platform services that power AI agent workflows — request/event ingestion, agent orchestration, document and file handling — improving reliability and delivery speed for product teams building on the platform.
- Contribute to the CI/CD and progressive-delivery infrastructure that lets both customer-facing product systems and internal developer tooling ship safely and continuously — including release-decoupling work that separates code deployment from customer-facing exposure.
- Build and maintain the cloud infrastructure the platform runs on — provisioning, capacity, and environment configuration — using Infrastructure-as-Code practices.
- Make sound architectural decisions for the services you own, with a focus on scalability, reliability, and maintainability within a microservices/event-driven system.
- Collaborate with product, applied-AI, and infrastructure teams on cross-functional initiatives, contributing to design and helping de-risk complex, multi-system projects.
- Uphold high engineering standards — code quality, automated testing, observability, CI/CD, and operational excellence — in your own work and in code review.
- Build and maintain observability into our services — metrics, logging, and tracing — that gives the team visibility into system health and speeds up incident diagnosis.
- Strengthen monitoring, incident response, and root-cause analysis for production systems running on managed cloud AI runtimes (e.g. AWS Bedrock AgentCore) and Kubernetes.
- Participate in an on‑call rotation to support our customer‑facing systems and critical internal tooling — treating every incident as an opportunity to protect the customer experience.
- Mentor junior and mid‑level engineers, sharing knowledge and supporting their technical growth.
About You
- 4+ years of professional software engineering experience, including experience designing, building, and operating backend systems in production.
- Strong proficiency in Python (FastAPI or similar) or another backend language, with experience in distributed systems, microservices, and cloud-native development.
- Hands‑on expertise with relational databases (PostgreSQL), caching/messaging systems (Redis), API design, and AWS, including container orchestration with Kubernetes (EKS).
- Experience with CI/CD pipeline engineering and progressive/safe delivery patterns (e.g. staged rollouts, traffic‑shifted or environment‑based release strategies).
- Infrastructure-as-Code practices for managing cloud infrastructure.
- Hands‑on experience with observability tooling (e.g. OpenTelemetry, Datadog) for building and operating production systems.
- Experience operating systems built on or integrating with LLM/Generative AI infrastructure (e.g. managed agent runtimes, LLM gateways, RAG or document‑retrieval pipelines) is a strong plus.
- Strong system design, debugging, and performance optimization skills, with the ability to make thoughtful trade‑offs between speed, quality, and long‑term scalability.
What’s in it For You?
- Hybrid Work Model: We’ve adopted a flexible hybrid working environment for our office‑based roles while delivering a seamless experience that is digitally and physically connected.
- Flexibility & Work‑Life Balance: Flex My Way is a set of supportive workplace policies designed to help manage personal and professional responsibilitie