Hugging Face Blog·· 2023-02-15AI 评分27
我们为何转向 Hugging Face Inference Endpoints,或许你也该考虑
Why we’re switching to Hugging Face Inference Endpoints, and maybe you should too
AI 导读
Hugging Face 推出托管服务 Inference Endpoints,可将 Hub 上的模型部署到 AWS 等云端的多种实例类型(含 GPU)。团队将原本跑在 AWS ECS + Fargate 上的 CPU 推理模型迁移过去,部署流程从六步简化为三步。测试显示 large 实例延迟约 80ms,比原方案快一倍以上,但成本高出 24% 至 50%。
来源:Hugging Face Blog · huggingface.co