Penn Interactive Logo

Penn Interactive

Senior Site Reliability Engineer

Posted 25 Days Ago
Remote
Hiring Remotely in Canada
Senior level
Remote
Hiring Remotely in Canada
Senior level
Own and operate infrastructure for a large-scale sports betting and media platform across cloud and production environments. Lead infrastructure migrations, build Kubernetes platform tooling and CI/CD automation, improve observability and incident response, support development teams, troubleshoot distributed systems, reduce operational toil, and mentor engineers.
The summary above was generated by AI

PENN Entertainment, Inc. is North America’s leading provider of integrated entertainment, sports content, and casino gaming experiences. From casinos and racetracks to online gaming, sports betting and entertainment content, we deliver the experiences people want, how and where they want them.

We’re always on the lookout for those who are passionate about creating and delivering cutting-edge online gaming and sports media products. Whether it’s through Hollywood Casino, theScore Bet Sportsbook, or theScore media app, we’re excited to push the boundaries of what’s possible. These state-of-the-art platforms are powered by proprietary in-house technology, a key component of PENN’s omnichannel gaming and entertainment strategy.

When you join PENN Entertainment’s digital team, you’ll not only work on these cutting-edge platforms through theScore and PENN Interactive, but you’ll also be part of a company that truly cares about your career growth. We’re committed to supporting you as you expand your skills and explore new opportunities.

With locations throughout North America, you can build a future at PENN Entertainment wherever you are. If you want to challenge conventions in gaming, media and entertainment, we want to talk to you.

About the Role & Team
The SRE team at PENN Entertainment is looking for a Senior Site Reliability Engineer to help build and operate the infrastructure behind a large-scale sports betting and media platform. You'll own critical infrastructure across compute, networking, storage, and cloud services (GCP/AWS) — driving complex migrations, building  platform tooling and automation (ArgoCD, Helm, GitHub Actions), and improving observability and incident response across hundreds of production services spanning multiple regulated jurisdictions. We're looking for someone with strong Kubernetes and distributed systems experience, proficiency in Go, Python, or Bash, and a track record of leading cross-team infrastructure projects with high autonomy. You'll solve ambiguous problems, reduce operational toil through automation, mentor teammates, and bring a production-first perspective to architecture decisions — all on a team that values pragmatic engineering, real ownership, and continuous improvement of how we work.

About the Work

  • Drive complex infrastructure migrations and projects — scoping, planning, execution, and validation across multiple production environments and jurisdictions
  • Build and maintain platform tooling and automation — ArgoCD, Helm, GitHub Actions, release pipelines, and service onboarding workflows that reduce toil for SRE and development teams
  • Support development teams — consult on infrastructure needs, unblock cross-team dependencies, review architecture proposals, and help teams adopt platform tooling and best practices
  • Design and improve observability and alerting — Datadog monitors, dashboards, and runbooks that surface meaningful signals and make systems operable by the whole team
  • Provide operational support and incident response — investigate and resolve production issues through structured debugging and root cause analysis, and contribute to on-call rotations to maintain platform reliability

About You

  • 5+ Years of Experience in a similar role (DevOps, Site Relatability Engineer)
  • Strong experience operating and troubleshooting Kubernetes in a production Linux environment (cluster lifecycle, networking, storage, scheduling)
  • Experience working with AWS, GCP, and/or on-premise environments
  • Proficiency in at least two of: Go, Python, Bash/Shell — for building tooling, automation, and debugging production systems
  • Deep understanding of distributed systems — failure modes, networking fundamentals, capacity planning, and performance analysis
  • Experience with GitOps and CI/CD workflows (ArgoCD, Helm, GitHub Actions, or similar)
  • Experience with infrastructure-as-code (Terraform, Helm, or equivalent)
  • Track record of leading complex migrations or infrastructure projects with cross-team dependencies
  • Strong incident response and troubleshooting skills — structured debugging across multiple services and infrastructure layers
  • Clear technical communication — you write docs others can use, explain trade-offs to varied audiences, and proactively unblock cross-team dependencies

Nice to have

  • Experience with service mesh technologies (Istio, Cilium)
  • Familiarity with distributed storage systems (Ceph, or similar)
  • Experience with bare-metal Kubernetes or Talos OS
  • Exposure to regulated environments (sports betting, fintech, or similar compliance-heavy domains)
  • Experience with Datadog or comparable observability platforms at scale
  • Familiarity with database operations — PostgreSQL, connection pooling (PgBouncer), or database migration tooling

What We Offer

  • Competitive compensation package.
  • Comprehensive Benefits package.
  • Fun, relaxed work environment.
  • Education and conference reimbursements.

#LI-REMOTE


Salary Range
$145,000$193,000 CAD

Penn Interactive is proud to be an equal opportunity workplace. We will consider all qualified applicants for employment without regard to race, color, religion, age, sex, sexual orientation, gender identity, national origin, disability, veteran status, genetic information, or any other basis protected by applicable law.Base pay is one part of the Total Rewards that Penn Interactive provides to compensate and recognize employees for their work. Most sales positions are eligible for a Commission under the terms of an applicable plan, while most non-sales positions are eligible for a Bonus. Additionally, Penn Interactive provides best-in-class benefits to eligible employees. We believe that benefits should connect you to the support you need when it matters most, and should help you care for those who matter most. That’s why we provide an array of options, expert guidance and always-on tools, that are personalized to meet the needs of your reality – to help support you physically, financially and emotionally through the big milestones and in your everyday life.

Similar Jobs

49 Minutes Ago
Remote
Canada
Senior level
Senior level
Software
Operate and maintain highly available, secure, containerized SaaS applications across AWS and Azure. Responsibilities include rotating 24x7 on-call coverage, observability, incident and security response, disaster recovery, Terraform infrastructure-as-code, CI/CD automation, cloud integration, performance optimization, and developing AI agents to automate SRE and DevSecOps workflows.
Top Skills: Amazon Web ServicesCheckovClaude CodeDatadogDockerDynatraceGithub ActionsGitlab CiGoJavaScriptKubernetesAzureNew RelicPrisma CloudPythonTerraformWiz
4 Days Ago
Remote
Canada
Senior level
Senior level
Cloud • Security • Software • Generative AI
Own and improve Elastic’s observability infrastructure across hosted cloud deployments. Responsibilities include Terraform-based infrastructure delivery, Python and Go development, production operations, incident response, on-call participation, RCA and postmortem writing, code and design reviews, mentoring, and improving operational documentation and processes. The role also involves operating Linux and containerized workloads, delivering complex projects independently, and maintaining secure, reliable platform infrastructure.
Top Skills: AnsibleArgocdBeatsElastic Cloud Enterprise (Ece)Elastic Cloud Hosted (Ech)Elastic Cloud On Kubernetes (Eck)ElasticsearchGoHelmKibanaKubernetesKyvernoLinuxLogstashPuppetPythonTeleportTerraformVault
6 Days Ago
Remote
Canada
Senior level
Senior level
Software • Automation
Owns reliability, scalability, observability, and incident response for a mission-critical SaaS platform. Responsibilities include 24x7 on-call support, root cause analysis, automation, AWS infrastructure design, EKS/Kubernetes and Docker operations, Terraform and Helm deployments, CI/CD maintenance, cloud networking, monitoring with Datadog, Grafana, and Prometheus, database operations, and migration toward Kubernetes. The role also develops self-healing systems, documentation, runbooks, and cross-functional customer-focused reliability practices.
Top Skills: AlbAmazon RdsAWSBashCi/CdDatadogDevsecopsDockerDocker SwarmEksGrafanaHelmIamInfrastructure As CodeKubernetesLinuxNlbPostgresPrometheusPythonRoute 53TerraformTerraformTransit GatewayVpcVpn

What you need to know about the Ottawa Tech Scene

The capital city of Canada and the nation's fourth-largest urban area, Ottawa has proven a rapidly growing global tech hub. With over 1,800 tech companies, many of which are leaders in their sectors, the city's tech talent now makes up more than 13 percent of its total workforce. This growth is driven not only by the big players like UL Solutions and Dropbox, but also by a thriving startup ecosystem, as new businesses emerge to follow in the footsteps of those that came before them.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account