作者训练 1.26M 参数模型把 htop、vim 等 TUI 转成真实 UI 组件
Instead of another GPU terminal renderer, I trained a 1.26M-param model to turn TUIs (htop, vim, emacs…) into real UI components [R]
作者训练了一个 1.26M 参数、5 MB 的轴向 Transformer,为终端每个单元格标注边框、标题、菜单项、选中行等 15 种角色,再由确定性代码把区域转成 A2UI 组件。
| Write-up + demo: https://drksci.com/labs-phosphene So this started as a bit of a gripe. Modern terminal renderers are seriously impressive and seriously complicated. GPU glyph atlases, texture caches, custom shaders, HarfBuzz shaping, ligatures, damage tracking, grid diffing, dirty-row uploads. Alacritty, Kitty, WezTerm and Ghostty are all doing heroic work to draw what is, at the end of the day, a grid of characters really fast. And every client still does the same thing at the end of it. Parse an escape-code stream, keep a cell grid, paint characters. Faithful, but opaque. Your phone can't reflow it, a screen reader gets a wall of box-drawing characters, and an agent has to squint at So I wondered: what if instead of throwing more GPU at drawing the grid, you used a bit of AI to understand it, once, server-side? Then send the client actual UI instead of a terminal.
Trained on public asciinema recordings. The labelling was done by Claude subagents, with a synthetic TUI generator for exact labels, all on a free-ish Colab T4. Honest numbers, because I'd rather say them before someone else does:
There's a replay with 8 apps (vim, htop, less, dialog, emacs, top, tig, nano). The native terminal sits on the left and the generated UI on the right, in sync, with every element outlined. Write-up + demo: https://drksci.com/labs-phosphene [link] [留言] |
来源:r/MachineLearning · reddit.com