B

Platform Engineer - (Site Reliability Engineering)

Bitso · 交易所 · 上架 2026-07-27
面议
运维/SRE/基础设施中级📍 Latin America远程

职位要求 / 描述

<div class="content-intro"><h3 class="p1"><span class="s1">Working At Bitso</span></h3> <p class="p1"><span class="s1" style="font-size: 12pt;">We are a diverse team that takes pride in understanding the perspectives of others. We fully embrace working remotely and we are eager to act, improve and accelerate progress inside and outside of our organization.</span></p> <p class="p1"><span class="s1" style="font-size: 12pt;">To drive revolutionary changes in society and make crypto useful, we delight our customers with world-class products, deep care, and intentional empathy.</span></p></div><h3><span style="font-size: 12pt;">Your Purpose </span></h3> <p><span style="font-size: 12pt;">At Bitso, reliability isn’t an afterthought — it’s a competitive advantage. As a Platform Engineer 2 focused on Incident Management, you’ll own the full incident lifecycle: from active response during live incidents, to driving postmortems, building automation, and eliminating the root causes that create toil in the first place. You’ll be the person who asks “how do we make sure this never happens again?” — and then actually builds it. If you thrive under pressure, love automation, and want to make a measurable dent in how a high-scale crypto platform operates, this role was designed for you.</span></p> <h3><span style="font-size: 12pt;">Reports To</span></h3> <p><span style="font-size: 12pt;"><strong>Incident Management Manager</strong></span></p> <h3><span style="font-size: 12pt;">Who You Are </span></h3> <ul> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Proven ability to operate confidently in high-pressure incident scenarios, including communicating clearly with senior stakeholders and leadership while a production issue is live</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Hands-on experience with Kubernetes — comfortable deploying, debugging, and navigating pod-level issues</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Solid understanding of CI/CD pipelines and modern DevOps practices</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Software development background in any language; ability to read, write, and debug code is essential (Python or Java experience is a plus)</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Strong automation mindset: you identify repetitive toil and your first instinct is to eliminate it, not absorb it</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Experience building or working with AI agents or LLM-based workflows is highly desirable </span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Strong interpersonal and written communication skills</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Self-directed learner who doesn’t need a fully defined path to start contributing</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Fintech or crypto industry background is a plus — familiarity with the domain vocabulary accelerates onboarding and incident triage</span></li> </ul> <h3><span style="font-size: 12pt;">What You Will Do </span></h3> <ul> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Own and execute on-call shifts end-to-end: acknowledge pages within SLA, declare incidents, assign roles, maintain comms cadence, and drive to resolution</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Build automation that drives the Sev1/Sev2 postmortem workflow — from scheduling and facilitation reminders to action-item assignment, ownership tracking, and due-date enforcement</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Leverage AI to identify patterns across incidents and propose systemic fixes: runbook improvements, alert tuning, platform hardening, and process changes</span></li> <li style="font-size: 12pt;"><span style="font-size: 12pt;">Build and extend internal automatio

技能关键字

#Python#Kubernetes#CI/CD#SLO/SLA#On-call#AI

职责方向

团队管理稳定性保障故障/值班监控告警部署发布自动化

数据来自公开渠道整理,薪资为公开 JD 或聚合估算,仅供参考,以面试谈薪为准。 ← 返回链聘 ChainHire 职位看板

Platform Engineer - (Site Reliability Engineering) · Bitso
面议
立即投递 →