Learning to Learn a Language:用合成非语言先验实现自然语言的上下文学习
Learning to Learn a Language: in-context learning of natural language from a synthetic non-linguistic prior [R]
论文提出一种语言先验,每条训练序列都来自随机采样的循环因果模型,即一种新的合成语言。一个仅在这类合成序列上训练的 300M 参数字节级 Transformer,在冻结权重下阅读维基百科文本时,下一字节预测随阅读量增加而变好,在英语、中文、印地语、阿拉伯语、日语、韩语六种语言上从 8 bits per byte 降至百万字节后的 0.9–2.4。
Learning from data as we observe it is easy for humans, but most machine learning models have limited ability to learn from new data that they have not seen during training. Prior-fitted networks (the idea behind TabPFN) showed that a model trained only on synthetic data can learn from real tabular data entirely in context.
I wanted to share our paper "Learning to Learn a Language" where we extend the idea to structured sequences such as natural language. We propose a prior over languages: every training sequence comes from a randomly sampled recurrent causal model, so each one is a new synthetic "language". A 300M-parameter byte-level transformer trained only on these synthetic sequences learns to predict real languages in context. Given Wikipedia text with frozen weights, its next-byte predictions get better the more it reads, in all six languages we tested (English, Chinese, Hindi, Arabic, Japanese, Korean), from 8 bits per byte down to 0.9–2.4 after a million bytes.
The same model also learns to count, to compare numbers, to add approximately, and to predict deterministic sequences such as the primes or the Kolakoski sequence, entirely in context.
It is of course still far worse on text than classical language models that are trained on trillions of tokens, while our model sees at most a million bytes of a language at test time. What we find interesting is that the ability to learn a language in context can come from a synthetic non-linguistic prior.
Paper: https://arxiv.org/abs/2610.05879
Code: https://github.com/cbl/prior-fitted-language-model
Weights: https://huggingface.co/lennartcb/pflm1
来源:r/MachineLearning · reddit.com