Jobiglo

No results.

SRE运维工程师

赛舵智能 · Shanghai

Mid zh
Linux Kubernetes Docker Prometheus Grafana Shell Python Go TCP/IP DNS HTTPS/TLS 负载均衡

Job description

职位概况

负责云平台及自建机房的基础设施运维,保障系统高可用、稳定运行,支持跨区域业务。需要在生产环境中独立处理故障、发布和容量管理。

主要职责

  • 维护阿里云、腾讯云、AWS、GCP、Azure 等公有云及自建机房的基础设施。
  • 负责 Kubernetes 集群及工作负载的日常运维、升级、扩缩容、网络存储和安全。
  • 建设并维护监控、日志、链路追踪和告警平台,完善 SLI/SLO、告警分级和故障定位。
  • 参与 7×24 On‑call,处理生产故障并完成根因分析与改进闭环。
  • 推动 CI/CD、配置管理和自动化,使用 Shell、Python、Go 编写工具脚本。
  • 进行容量、性能和成本分析,落地高可用、容灾和成本优化方案。
  • 与研发、测试、安全和业务团队协作,完善发布、变更和应急流程。
  • 沉淀架构图、Runbook、故障案例和运维标准,提高团队效率。

任职要求

  • 3 年以上生产环境运维、发布或故障处理经验。
  • 熟悉 Linux 系统及常用网络诊断工具,能定位 CPU、内存、磁盘 IO、网络性能问题。
  • 具备 Kubernetes 与 Docker 实践经验,能独立处理 Pod 生命周期、调度、服务发现等。
  • 熟悉 Prometheus、Grafana 或同类可观测平台,能设计指标、告警规则和 Dashboard。
  • 了解日志、链路追踪平台的使用与运维。
  • 掌握 TCP/IP、DNS、HTTPS/TLS、反向代理和负载均衡原理。
  • 能够使用 Shell、Python 或 Go 开发自动化脚本,关注幂等性和错误处理。
  • 具备高可用、容量管理、备份恢复或容灾实践,了解 RTO/RPO 等指标。
  • 接受合理排班的 7×24 On‑call,具备清晰沟通和风险判断能力。

必备技能

  • Linux
  • Kubernetes
  • Docker
  • Prometheus
  • Grafana
  • Shell
  • Python
  • Go
  • TCP/IP
  • DNS
  • HTTPS/TLS
  • 负载均衡

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec 赛舵智能.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

A question about this job?

Ask it here: you will get the full job summary by e-mail, right away.

💬 Chat with us on Telegram

Published 1 month ago

Expires 1 week from now

49 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

赛舵智能

Shanghai