跳到正文
r/MachineLearning· /u/Correct_Positive_108·· 2 小时前AI 评分29

开发者吐槽 Together AI 限流:跑 Llama 3.3 70B 和 Qwen 2.5 多智能体测试时撞上 RPM/TPM 上限

Looking for developer-friendly inference providers who give you enough API credits to experiment [D]

AI 导读

一名独立开发者反映,在使用 Together AI 并行运行多个智能体、调用 Llama 3.3 70B 和 Qwen 2.5 做仓库索引与基准生成时,很快触发了 RPM/TPM 限流,导致无法进行有意义的测试。他表示模型本身没问题,但作为没有企业级收入的个人开发者,难以承担升级成本,因此想寻找能提供足够 API 额度供实验的开发者友好型推理服务商。

正文

I’m hitting rate limits on Together AI. For context, I’ve been working on an agentic repository indexing and benchmark generation tool, and I’m running multiple agents in parallel across models like Llama 3.3 70B and Qwen 2.5.

When I first started working on this, Together AI was great. But once I graduated from toy scripts to running multiple agents, I started running into RPM/TPM limits pretty quickly. The annoying part is that the models themselves are fine. I just can’t actually run enough requests at once to do meaningful testing.

Yes I know I could upgrade but I’m a solo dev. I don’t have enterprise level revenue. Maybe someday lol but not yet.

submitted by /u/Correct_Positive_108
[link] [留言]

来源:r/MachineLearning · reddit.com