skypilot-multi-cloud-orchestration

仓库创建 2026年3月17日最近提交 2 个月前SkillHot 收录 21 天前
▸ 精选理由

便于在多云/现货实例上最低成本运行训练与批处理

▸ 风险提示

会在云上创建实例并产生费用,需注意账号权限与账单

这个 Skill 做什么

跨云自动调度 ML 作业并优化成本与抢占实例恢复策略。

在多云环境(如 AWS/GCP/Azure)上调度训练或批处理作业,自动选区、优化成本并智能使用 spot/抢占实例与自动恢复策略。适合要跨云跑长期训练或想压低 GPU 成本的场景。特点是自动择优云/区域和抢占实例恢复,能最大化性价比并减少人工干预。

▸ 展开 SKILL.md 英文原文

Multi-cloud orchestration for ML workloads with automatic cost optimization. Use when you need to run training or batch jobs across multiple clouds, leverage spot instances with auto-recovery, or optimize GPU costs across providers.

自动化集成多云编排成本优化GPU调度通用
1.5k
Stars
104
Forks
16
仓库内 Skill
+0
7 日增星
安装 / 使用
给你的 Agent 一句话(通用)
帮我安装这个 skill:https://raw.githubusercontent.com/OpenRaiser/NanoResearch/main/skills/vendor-ai-research/skypilot/SKILL.md
或 curl 直取 SKILL.md
curl -fsSL "https://raw.githubusercontent.com/OpenRaiser/NanoResearch/main/skills/vendor-ai-research/skypilot/SKILL.md"
SKILL.MD 节选查看完整文件 ↗
# SkyPilot Multi-Cloud Orchestration

Comprehensive guide to running ML workloads across clouds with automatic cost optimization using SkyPilot.

## When to use SkyPilot

**Use SkyPilot when:**
- Running ML workloads across multiple clouds (AWS, GCP, Azure, etc.)
- Need cost optimization with automatic cloud/region selection
- Running long jobs on spot instances with auto-recovery
- Managing distributed multi-node training
- Want unified interface for 20+ cloud providers
- Need to avoid vendor lock-in

**Key features:**
- **Multi-cloud**: AWS, GCP, Azure, Kubernetes, Lambda, RunPod, 20+ providers
- **Cost optimization**: Automatic cheapest cloud/region selection
- **Spot instances**: 3-6x co
via SKILL·HOT · 数据来自 GitHub 公开信息 · 原文版权归作者所有