跳到正文
r/MachineLearning· /u/neuralbeans·· 3 小时前AI 评分34

在 API 模型上跑 benchmark 时如何防止数据泄露?[D]

Best practices when running a benchmark on online models [D]

AI 导读

有开发者提出,为低资源语言构建 benchmark 时,本地模型无泄露风险,但仅通过 API 访问的模型可能将输入数据用于训练。他询问是否存在评估在线模型而不丢失输入数据的成熟方法,以及是否该信任 Google 和 OpenAI 关于付费账户输入不用于训练的说法。

正文

I'm developing a benchmark for a low resource language and I don't want it to be leaked and used for training when it is being used to get predictions. For locally run models it shouldn't be a problem, but for models that are only accessible via API, it is. Is there an established way to evaluate online models without the input data being lost? Do you trust Google and OpenAI when they say that they do not use your inputs for training when you have a paid account?

submitted by /u/neuralbeans
[link] [留言]

来源:r/MachineLearning · reddit.com