Job Description
About the Role
\n
Remitian is looking for a high-agency Site Reliability Engineer to own the reliability of the platform that moves money and tax data for our customers. Our systems run on a daily clock: payment files are prepared, submitted to banking partners, and reconciled against receipts on a fixed schedule every business day. When something in that chain is late or wrong, customers and tax authorities feel it the same day.
\n
You will own the infrastructure, observability, and automation that keep those flows dependable, and you will take your turn in our production monitoring rotation. That rotation is deliberately part of the job: we want an engineer who lives with the operational reality of the platform and then engineers the manual work out of it.
\n
This is a build-and-operate role, not a ticket queue. You will own meaningful parts of our cloud footprint, deployment pipelines, and monitoring, and you will have a direct hand in shaping how Remitian runs software in production.
\n
This role is based in either our Ottawa, Ontario office or in Florida, with Miami preferred, on a hybrid schedule of 3 to 4 days per week in office.
\n
Key Responsibilities
\n
Infrastructure and Cloud Platform
\n
- \n
- Own and evolve our cloud infrastructure, defined as code and deployed repeatedly across environments
- Build and maintain CI/CD pipelines that get changes to production safely and often
- Manage environment configuration, secrets, and access in a way that is auditable and least-privilege by default
- Improve deployment mechanics so releases are routine, low-drama, and reversible
- Manage networking, connectivity, and secure remote access for engineers and services
\n
\n
\n
\n
\n
\n
Reliability and Observability
\n
- \n
- Own monitoring, dashboards, logging, and alerting for the scheduled jobs and services that move money and data
- Make alerts actionable: reduce noise, add the context an on-call engineer actually needs, and retire alarms that no longer earn their place
- Define and track service level objectives for the workflows customers depend on
- Lead incident response, drive issues to resolution, and turn postmortems into completed follow-up work
- Improve error handling, retries, and recovery paths for time-sensitive batch and integration workflows
\n
\n
\n
\n
\n
\n
Production Monitoring Rotation and Automation
\n
- \n
- Take your turn in the production monitoring rotation, verifying that daily payment and reconciliation workflows complete as expected and investigating when they do not
- Triage production alerts, diagnose root causes across services, jobs, queues, and third-party integrations, and communicate status clearly to engineering, support, and compliance stakeholders
- Treat every manual check as a defect to be automated: replace human verification with automated validation, self-healing jobs, and tooling that surfaces problems before a person has to look for them
- Build and improve internal tooling and dashboards that make the rotation faster, safer, and eventually unnecessary
- Help train and onboard other engineers into the rotation, and keep the runbooks current
\n
\n
\n
\n
\n
\n
Security and Compliance Support
\n
- \n
- Implement and maintain access controls, audit trails, and infrastructure guardrails appropriate to a regulated fintech
- Keep dependencies, images, and infrastructure patched and current
- Support audit and compliance requirements with evidence, documentation, and tooling rather than manual effort
- Handle production data with care, applying the principle of least privilege to yourself as well as to everyone else
\n
\n
\n
\n
\n
Ownership and Continuous Improvement
\n
- \n
- Take real ownership of infrastructure, tooling, and operational workflows
- Identify problems, propose solutions, and drive them to completion without waiting for direction
- Push back, raise concerns, and surface ideas that improve reliability or how the team works
- Contribute to engineering standards, technical documentation, and onboarding
\n
\n
\n
\n
\n
Who You Are
\n
- \n
- Three to seven years of experience in site reliability, platform, or infrastructure engineering
- Strong hands-on experience with a major cloud provider, ideally AWS, including compute, networking, storage, and identity
- Comfortable with infrastructure as code and with CI/CD pipelines as a first-class part of your work
- Comfortable writing code, not just configuration, with scripting and automation in Python, Bash, C#, TypeScript, or similar
- Practical experience with observability tooling, including logs, metrics, dashboards, and alerting
- Experience being on call for production systems, and a track record of reducing operational load rather than absorbing it
- Strong debugging instincts across distributed systems, scheduled jobs, queues, and third-party integrations
- High agency, meaning you find work that needs doing, take ownership, and move things forward without being told
- Strong written and verbal communicator, able to explain incidents, trade-offs, and risk clearly to technical and non-technical audiences
- Sound judgment around production access, sensitive data, and change management
- Excited to work in a fast-paced startup environment where priorities shift and ambiguity is part of the job
- Based in or willing to relocate to either the Ottawa, Ontario area or Florida, with Miami preferred, and excited to be in the office 3 to 4 days per week
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
Nice to Haves
\n
- \n
- Experience in fintech, payments, banking, or another regulated industry
- Exposure to batch payment file processing, bank integrations, or financial reconciliation workflows
- Experience operating .NET workloads in production
- Experience with containers and managed container orchestration services
- Experience with zero-trust or modern remote access tooling
- Experience building internal tools, dashboards, or developer tooling
- Experience using LLMs and AI tooling to automate operational work or accelerate development
- Experience in a startup where you wore multiple hats and owned production end to end
\n
\n
\n
\n
\n
\n
\n
\n
\n
Why Join Us
\n
- \n
- Own the reliability of systems that move real money on a daily deadline, where your work is visible and matters immediately
- Join a high-performing team building mission-critical fintech infrastructure
- A clear mandate to automate, since we are actively investing in replacing manual operational work with tooling, and you will lead that effort
- Broad scope across cloud infrastructure, deployment, observability, security, and internal tooling
- Use modern tools, including AI and LLM tooling, to move quickly and build with leverage
- Hybrid work from either our Ottawa or Florida location, where the team comes together 3 to 4 days per week to build, ship, and learn alongside one another
\n
\n
\n
\n
\n
\n
\n
