跳到正文
r/MachineLearning· /u/BuckChancey·· 3 小时前AI 评分52

作者训练 1.26M 参数模型把 htop、vim 等 TUI 转成真实 UI 组件

Instead of another GPU terminal renderer, I trained a 1.26M-param model to turn TUIs (htop, vim, emacs…) into real UI components [R]

AI 导读

作者训练了一个 1.26M 参数、5 MB 的轴向 Transformer,为终端每个单元格标注边框、标题、菜单项、选中行等 15 种角色,再由确定性代码把区域转成 A2UI 组件。

正文
Instead of another GPU terminal renderer, I trained a 1.26M-param model to turn TUIs (htop, vim, emacs…) into real UI components [R]

Write-up + demo: https://drksci.com/labs-phosphene

So this started as a bit of a gripe. Modern terminal renderers are seriously impressive and seriously complicated. GPU glyph atlases, texture caches, custom shaders, HarfBuzz shaping, ligatures, damage tracking, grid diffing, dirty-row uploads. Alacritty, Kitty, WezTerm and Ghostty are all doing heroic work to draw what is, at the end of the day, a grid of characters really fast.

And every client still does the same thing at the end of it. Parse an escape-code stream, keep a cell grid, paint characters. Faithful, but opaque. Your phone can't reflow it, a screen reader gets a wall of box-drawing characters, and an agent has to squint at │ ▶ item │ to work out which row is selected.

So I wondered: what if instead of throwing more GPU at drawing the grid, you used a bit of AI to understand it, once, server-side? Then send the client actual UI instead of a terminal.

  • a tiny model (1.26M params, 5 MB, an axial transformer over rows and columns) labels every cell with a role: border, title, menu item, selected row, table, input, status bar, key hint, etc. (15 roles)
  • dumb deterministic code turns those regions into A2UI components (Google's declarative UI stream protocol): lists, text fields, buttons, progress bars
  • once a screen layout has been seen, it locks as a template and only the changed content goes over the wire as JSON-pointer patches. The model doesn't even run.
  • pressing a button in the UI sends the keystroke back. F10 is just a Button with an action.

https://preview.redd.it/srgc7o4b06uh1.png?width=2100&format=png&auto=webp&s=a390b467df08095c6941643dc0cd5b1b1180a91b

Trained on public asciinema recordings. The labelling was done by Claude subagents, with a synthetic TUI generator for exact labels, all on a free-ish Colab T4.

Honest numbers, because I'd rather say them before someone else does:

  • accuracy on held-out real screens is mIoU 0.51. Usable, not amazing; it's a first labelling round of 600 frames.
  • 40% of ~14k screens never touch the model (template hit). On less and dialog it's ~90%, on htop and nano it's rubbish because the meters keep changing the layout.
  • The A2UI stream is ~25× bigger than raw VT. VT is a stupidly compact format, turns out. The win is the client never runs a terminal emulator at all, not bandwidth.

There's a replay with 8 apps (vim, htop, less, dialog, emacs, top, tig, nano). The native terminal sits on the left and the generated UI on the right, in sync, with every element outlined.

Write-up + demo: https://drksci.com/labs-phosphene
Code, labels, results: https://github.com/drksci/phosphene

submitted by /u/BuckChancey
[link] [留言]

来源:r/MachineLearning · reddit.com