Andrew Murphy

Sr. Site Reliability Engineer

Site Reliability Engineer who builds secure, resilient, cost-efficient systems at scale. I work at the intersection of infrastructure, automation, and developer experience, focused on making things genuinely better for the people who use them, not just the dashboards. Comfortable setting technical direction and just as comfortable rolling up my sleeves to build it.

Work

Sr. Cloud Operations Engineer

Led reliability initiatives for Autodesk's cloud-based products in the media and entertainment industry, combining hands-on systems engineering with strategic architectural support.

  • Co-led replacement of an end-of-life routing layer carrying all customer traffic for a SaaS platform serving ~16,000 tenant sites. Authored the architecture decision records and building the automated configuration service that replaced manual routing management
  • Delivered multi region SaaS Platform expansion enabling data-residency compliance and latency optimizations for European and Asia Pacific customers, building and validating the live customer site migration workflow, beta program, and publishing throughput benchmarks that made migration timelines predicable at scale
  • Migrated production observability for a business-critical platform off two sunsetting vendor tools, leading the gap analysis that mapped legacy metrics and alerts to replacement equivalents and retiring obsolete monitoring credentials and dashboards
  • Automated provisioning, teardown, and recovey workflows accross a fleet of single tenantcy customer environments, optimizing for cost, performance, and reliability while reducing toil and human error
  • Established the operational readiness standards governing how the team assumes on-call ownership of new services, defining alert quality gates, escalation rules, runbook requirements, and mandatory knowledge transfer sessions
  • Authored 13 production incident postmortems accross 6 years and modernized 11 service runbooks, including operational documentation review for partner engineer teams
  • Quantified 2,400 in monthly avoidable logging spend accross 15,900 tenants and automated remediation, with modeleted annual savings of approximately $29,000

Systems Engineer

Owned infrastructure and platform reliability for core applications including EthicsPoint and NAVEX Gateway (IDP), with a focus on automation, observability, and security posture.

  • Automated provisioning and configuration of 1,000+ Windows and Linux servers, load balancers, and firewalls across hybrid environments, replacing manual infrastructure workflows and improving deployment consistency and operational reliability
  • Supported high-availability production systems by building and maintaining infrastructure monitoring and critical workflow validation across New Relic and Check_MK, improving visibility into service health and reducing operational risk
  • Improved disaster recovery readiness by developing cross-region replication workflows and comprehensive validation test plans, establishing repeatable processes for verifying recovery capabilities
  • Led application release coordination and outage planning across DevOps, QA, and product teams, using Octopus Deploy to standardize deployments and coordinate production changes across critical services
  • Drove improvements to the organization’s security posture through evaluation and implementation of Cortex XDR, Imperva WAF, BigFix, and Prisma Cloud, strengthening endpoint, application, patching, and cloud security controls
  • Provided 24x7 operational support for business-critical applications through an on-call rotation, leading incident response, troubleshooting production issues, and coordinating remediation to maintain platform availability

Network/Operations Engineer

Supported the development and operations of a distributed compute platform for 3D rendering, with a focus on performance, reliability, and customer tooling.

  • Improved GPU farm efficiency by 4% through infrastructure tuning, system optimization
  • Diagnosed and resolved failures across rendering infrastructure, restoring degraded systems to service and maintaining GPU farm availability
  • Provided technical support for customers using the rendering platform, troubleshooting application, infrastructure, and workflow issues while identifying recurring pain points and opportunities for automation
  • Developed a Python-based Cinema 4D plugin informed by recurring customer support issues, streamlining common rendering workflows and reducing manual troubleshooting and support effort
  • Represented the company at NAB Las Vegas and SIGGRAPH Los Angeles, engaging with customers and prospects at exhibitor booths, demonstrating products, answering technical questions, and supporting sales conversations

Junior Systems Administrator Intern

Developed skills in systems and network administration, virtualization, and cloud computing in an MSP and IaaS environment

  • Expanded technical expertise across enterprise infrastructure and systems administration, working directly with systems administrators to support and troubleshoot platforms including EM7, Confluence, VMware, and vSphere
  • Led the planning, implementation, and migration of internal documentation and knowledge assets to a new Atlassian Confluence environment, improving organization and accessibility of company-wide technical documentation

Projects

Computehub by NWHTC

Started as a cryptocurrency mining operation, pivoted to a low-cost leased cloud computing platform; wound down in 2024

Built withCPU / GPU / Cloud Computing

Education

Skills

Cloud & Infrastructure
AWS / Linux / Windows Server / Ephemeral Architecture / Auto Scaling / Load Balancers / Web Servers / CDN / WAF / Serverless / Firewalls / Switching / Hypervisors
Infrastructure as Code & Automation
Terraform / AWS CDK / CloudFormation / Jenkins / GitHub Actions / GitLab CI / Puppet / CI/CD Pipelines / Configuration Management
Containers & Orchestration
Docker / ECS
Programming & Scripting
Python / Bash / Ruby / JavaScript / Postgres
Observability & Monitoring
New Relic / Dynatrace / check_mk / ELK Stack / Splunk / Grafana / Prometheus
Incident Management & Reliability
PagerDuty / BigPanda / ServiceNow / 24x7 On-Call / Incident Command / Root Cause Analysis / Postmortems / Runbooks / High Availability / Disaster Recovery
Security & Compliance
Cortex XDR / BigFix / Prisma Cloud / Snyk / Rapid7 / Vulnerability Management
Collaboration & Leadership
Incident Response Leadership / Architecture Reviews / Mentorship / Cross-Team Collaboration / Technical Strategy / Developer Enablement
Version Control & Workflow
Git / GitHub / GitLab / Branching Strategies / Code Reviews

Languages

English
Native speaker

Interests

Running
Hiking
Skiing
Traveling
Cooking

References

Available Upon Request