Terminal image passthrough
(feature: terminal-images)
Programs running inside a Terminal pane can draw pictures. TerminalScreen reads the Kitty graphics protocol out of the PTY byte stream, decodes it, and the renderer paints the result over the pane's text.
tui-lipan = { version = "0.18", features = ["terminal-images"] }cargo run --example terminal_images --features terminal-imagesIt does not depend on the host terminal
The child's escapes are never forwarded to the host. They are decoded into pixels and re-encoded through the same path the Image widget uses, so the host renders them with whatever it supports — Kitty, iTerm2, sixel, or half-blocks. A pane in a plain xterm shows pictures.
Decoding rather than forwarding is also what makes the rest work: image ids from two panes cannot collide, a pane that is half scrolled off gets its pixels cropped instead of squashed, and the cell-diff render pipeline is untouched.
Setting the cell size
A terminal deals in cells; a program drawing a picture needs pixels. It learns the conversion from the PTY's TIOCGWINSZ pixel fields or by asking with CSI 14 t. Give both ends the same answer:
use tui_lipan::prelude::*;
let cell = host_cell_size();
let mut screen = TerminalScreen::new(rows, cols, scrollback);
screen.set_cell_size(cell);
let config = TerminalPtyConfig::default().cell_size(cell);ManagedTerminal does this for you. A raw TerminalScreen defaults to a 10x20 cell, which is a guess: a mismatch shows up as images that overlap the text below them or leave a gap, because the child reserved a different number of rows than the pane drew.
Pass the cell size to later resizes too, with TerminalPty::resize_with_cell_size; plain resize keeps the last one it was given.
Where images live
Two ways an image gets placed
At the cursor (a=T without U=1) is what icat and friends do: the image lands where the cursor is, the cursor moves past it, and the placement is anchored to that scrollback line.
Through placeholder cells (a=T,U=1) is what terminal UI toolkits do, ratatui-image among them: the transmission draws nothing, and the program then writes the placeholder character U+10EEEE into the cells the image should cover, tagging them with the image id (in the cell's foreground colour) and the position inside the image (in combining marks). Those placements are read back off the grid on every snapshot rather than stored, so they scroll, clear, and reflow with the cells holding them, for free. A cell may leave any of that out and inherit it from its left-hand neighbour — including the high byte of the image id — which is what keeps a row of placeholders down to a single escape sequence.
Cursor placements
A cursor placement is anchored to an absolute scrollback line, in the same space as OSC 133 semantic marks. That is what makes it behave like the text it was drawn against:
| Event | What happens |
|---|---|
| Output scrolls | The image scrolls with it, cropped row by row as it leaves the viewport |
| Scrolling back | It reappears at the line it was drawn on |
| A line falls out of scrollback | The image's remaining rows stay; it goes once its last row is evicted |
Screen erased (ED 2) | The image scrolls into history with the text above it, as the screen does |
Scrollback erased (ED 3, as clear sends) | Images in the erased history go with it |
| Alternate screen | Placements made there are dropped when the child leaves it |
| Column resize | Kept, unless the change actually rewraps text — then the anchor stops naming what it named, and placements are dropped |
RIS / TerminalScreen::reset | Everything is cleared |
Each event applies to the placements that exist when it happens, so the outcome does not depend on how the stream was split into process_bytes calls. clear followed by an image leaves the image on screen whether the two arrive in one write or two, which matters to a multiplexer that relays a pane's output in larger chunks than the program wrote it.
A placeholder placement needs none of that bookkeeping: it is the text, so it does whatever the cells do, and it is gone the moment they are.
Each placement carries the image_id its transmission used. A renderer must key its encoding on that and not on the pixels alone: a host drawing through Kitty identifies a placement by the id of its encoding, so two placements sharing one encoding are a single placement to it — and two copies of one picture would collapse, the second silently not drawn.
TerminalRenderSnapshot::images carries the placements overlapping the visible rows, back to front by Kitty z-index. Rows and columns are viewport-relative, and a cursor placement's may be negative when the image starts above or to the left of the pane.
A transmission may name the cell box it is meant to fill (c= columns, r= rows), and a virtual placement is then mapped through that box rather than through the cell grid: each placeholder cell takes the slice of the image its position in the box calls for. This is what lets a child draw at whatever resolution it likes — twice the cell size on a HiDPI screen, say — and still have the picture scaled into the cells it asked for instead of cropped to its top-left corner. The box belongs to the image id, so a sender animating by re-transmitting under the same id declares it once.
Memory
Decoded pixels are capped per screen, 96 MiB by default, and evicted least-recently-used once the cap is passed — placed images included, since leaving old plots on screen must not pin memory. One image larger than the whole budget is kept anyway; that is a budget set too low, not a picture that should silently fail to appear.
screen.set_image_budget(16 * 1024 * 1024);Server-side terminal mirrors that forward the original PTY stream to another rendering client can avoid decoding every image twice:
screen.set_image_storage_enabled(false);That mode still validates command metadata, answers protocol queries, and applies image-implied cursor movement. It retains dimensions only; render snapshots contain no images. Raw RGB/RGBA payloads also skip base64 decoding, while PNG payloads are still decoded far enough to obtain their dimensions.
screen.has_images() is whether any of those pixels are still retained: true after a child has transmitted an image that has not been deleted or evicted, including when it has scrolled out of view. Visible placements this frame are snapshot.images. A host that needs the class of pane whose image layer does not follow a widget shrink or fade should look at the flag, not at the process in front of the PTY.
Payloads are bounded before decoding as well: 32 MiB per transmission, 16384 pixels per axis.
Out-of-band payloads
A child redrawing a full window every frame can leave its pixels somewhere and name them in the escape sequence instead of base64-encoding several megabytes through the PTY: t=f a file it leaves in place, t=t a temporary file, t=s a POSIX shared-memory object. This is the difference between a hundred bytes on the wire and a compress/encode/decode/inflate round trip per frame, and it is what lets a program in a pane animate at the host's own frame rate.
The last two are consumed by reading - the file is deleted, the object unlinked - so exactly one reader may ever claim them. Anything that fans one PTY stream out to several readers (a multiplexer with several attached clients, a server-side mirror) must therefore decline them and keep only t=f, which stays readable:
screen.set_image_media_policy(GraphicsMediaPolicy::SHARED);GraphicsMediaPolicy::NONE declines all three, so every picture arrives inline. Refusals are the protocol's own EBADF/ENOTSUPP reports, which is what makes a child fall back to t=d rather than lose the image. A t=t path must be temporary or say in its own name what it is for, t=s names are resolved as shared-memory objects rather than paths, and both reads are capped by the same transmission budget as an inline payload.
A path only means something on the machine that wrote it, so a pane attached from elsewhere should decline: a remote client reading /tmp/... from its own filesystem is the one failure mode here that is not a clean error.
Out the other side, to the host
The same reasoning applies to what this framework writes to the terminal it is running in, and the cost is the same one twice over: a pane whose child sends a full window of pixels every frame would otherwise have them deflated, base64-encoded, chunked and written down stdout, and the host would undo all of it.
So a frame goes into a POSIX shared-memory object and the escape sequence carries its name. Nothing configures this. At startup the host is asked - with a t=s query it can only answer by reading a real object - and only a terminal that answers OK is handed frames that way; everything else keeps the inline path, including a terminal reached through tmux, whose reader is not the terminal. Objects are unlinked by the host as it reads them, and by this process for any frame the host was never told about.
The question is only put to a host that could answer it, and that is judged from the environment before anything is written. A terminal on the other end of an ssh connection is never asked: it cannot resolve a name in this machine's memory whatever protocol it speaks, and TERM survives the hop, so it would otherwise look local and capable. Neither is a terminal that is not recognizably one of the implementations - kitty, Ghostty, WezTerm and Konsole, by TERM, TERM_PROGRAM, or the variable each one sets.
That matters because the query is an APC string sent before the alternate screen is entered. A terminal is supposed to consume an APC it does not implement, and most do; the ones that do not print it as text, and the mess lands in the scrollback and stays there. Not asking is cheap - an unasked host keeps the inline path, which is what an unsupported host does anyway.
For a host judged wrongly, TUI_LIPAN_GRAPHICS_SHM settles it: 1 asks a terminal that implements the protocol under a TERM this does not know, and 0 never asks, for a local terminal that mangles the question. The host must still answer OK either way - the variable controls the asking, not the result.
On Linux the objects are pooled, which is why /dev/shm holds a handful of tui-lipan-pool-<pid>-* names for as long as a pane is drawing. A fresh object per frame is a fresh allocation per frame - every byte written lands on a page that does not exist yet - so each slot is created once and written again every frame, and what the host is handed is a fresh hard link to it. That link is the part the host unlinks, and its disappearance is the only signal that the host is finished: a slot whose link is still there is never reused, and a frame that finds every slot busy allocates its own object exactly as it used to. Pool objects left behind by a run that was killed are removed by the next run, which can tell because the name carries the process that made it.
Windows has no POSIX shared-memory namespace, so frames there are always inline.
What is supported
| Key | Support |
|---|---|
a=t, a=T, a=p, a=d, a=q | Transmit, transmit-and-display, display, delete, query |
t=d | Direct transmission, chunked with m=1 |
| Payload | Base64, with or without = padding (Kitty's icat sends it unpadded) |
t=f, t=t, t=s | Out-of-band transmission: a file left in place, a temporary file, a POSIX shared-memory object |
f=24, f=32, f=100 | RGB, RGBA, PNG |
o=z | zlib-compressed payloads |
s, v | Pixel dimensions for the raw formats |
i, I, p | Image id, image number, placement id |
x, y, w, h | Source rectangle to display |
c, r | Explicit placement size in cells |
z | Stacking order against the text layer |
C=1 | Leave the cursor where it is |
q=1, q=2 | Suppress success reports / all reports; a later chunk's q= applies to the rest of its run |
d= | a/A, i/I, n/N, c/C, z/Z, p/P, x/X, y/Y |
U=1 | Virtual placements shown through Unicode placeholder cells |
Not supported, and answered with the protocol's own ENOTSUPP report so a child that probes first gets a clean answer rather than silence:
- The protocol's animation frames (
a=a,a=f,a=c). A sender that animates by re-transmitting under the same image id — which is whatratatui-imagedoes for GIFs — works regardless: the new pixels replace the old and the placeholders keep pointing at them. - Relative placements (parent references).
A payload that is not base64 is answered with EINVAL, under the command's id and q= like any other report, and nothing in it is acted on. A bad chunk, or one that takes a run past the size cap (EFBIG), abandons the whole m=1 run it belongs to: the error goes out once, under the run's id, and the run's remaining chunks up to its m=0 are consumed without a reply.
Sixel input from the child is a separate protocol and is not read.
Under a modal backdrop
A root-portal overlay's backdrop (Modal::backdrop_style, CommandPalette::backdrop_style) recolors the cells behind it. An image is not cells, and the host has no per-image opacity, so the renderer recolors the image's pixels instead, before encoding them: the backdrop's fill, dim_by, tint_by, and background transform, applied exactly as they apply to a cell background, so a picture dims to the same color as the cells beside it. Alpha is kept. Foreground transforms do not apply, because a picture reads as surface, not as text.
Layers inside the tree that recolor cells reach the pixels too, in the order the renderer applies them. Each runs the renderer's own cell code on the pixel's color, so the match is exact:
- An
EffectScopewhose effects depend only on a cell's own color:dim_by,lighten_by,tint_by,transform_bg,VisualEffect::Monochrome, andVisualEffect::PaletteQuantize, including underChannelsand a rectangularClipped.EffectScope::cells_only()opts a scope out: its effects recolor the cells and leave every image under it at its own pixels. - An
Animatedopacity toward a color,opacity_target, unless it isopacity_fg_only. - A
CanvasorCenterstacked over the image whose style dims, tints, or transforms without painting a background. That is how aLocal-scope modal's backdrop draws. Its dialog hides the image under it, as a root dialog does.
A layer reaches an image only if it applies after the image draws: an EffectScope or Animated around it, or a surface stacked above it. A layer inside an overlay reaches only the images in that overlay.
- Only covered cells dim. Root backdrops span the whole viewport, so in practice that is the whole visible image; the part under the dialog itself is hidden as before. A layer in the tree dims the cells it covers after its clip.
- One encode per image on open. The dimmed pixels are cached as their own variant beside the undimmed ones, and a Kitty stream transmits them under a separate image id. Closing the dialog over pixels that have not changed switches back to the undimmed image the host already holds. Pixels that changed while it was open, such as a live plot, have only dimmed encodes, so closing encodes the current frame undimmed.
- Never the wrong dimming. While an image re-encodes, its previous frame normally stays on screen. When the only previous frame is dimmed the other way, because the dialog just opened or closed, the renderer encodes the new frame during that paint instead, for every protocol. That costs one synchronous encode per image at each open or close, and none on later frames.
- Live pictures pay per frame. A picture that keeps changing under a dialog, such as a browser or a video, is recolored on every new frame. Dims, tints, and fills cost a few milliseconds for a full 1080p frame.
transform_bg(ColorTransform::Elevate(_)),Monochrome, andPaletteQuantizemix a color's channels, so a photographic frame under them costs more; past the first 16,384 colors, each is computed from the nearest color at 64 levels per channel, at most two levels away.cargo bench --bench image_backdrop --features terminal-imagesmeasures both. - Full strength at once. Images take the backdrop at its full strength from its first frame, and an
Animatedfade at the opacity it ends at, from the frame it starts. Following a fade would re-encode every image under it on every frame of the fade. - Half blocks are left to the backdrop. They are cells, which the backdrop already recolors.
- Kitty image ids survive. A placeholder cell that carries its image id in its foreground keeps that foreground under a backdrop, an
EffectScope, anAnimatedfade, or a surface in a layer that draws images; only the pixels dim.
Known limits
- Some layers leave the pixels alone. Effects that depend on where a cell is or on time (
Scanlines,Gradient,RainbowWave,RetroCrt,Ripple),ContrastPolicy, custom effects, mask clips, and anAnimatedopacity withoutopacity_targetrecolor cells without reaching the pixels under them. So does anEffectScopethat wraps an overlay's portal rather than sitting inside the overlay. - A painted surface covers an image. A
CanvasorCenterwhose style paints a background, even a translucent one, replaces the cells it covers, so the image stops showing there instead of dimming. - Reattaching to a session loses images drawn before the attach.
export_replay_bytesis a text replay stream and does not re-emit image payloads. - A partly visible image is cropped, not scaled. That is what makes scrolling look right, but it also means an image wider than its pane shows its left part rather than shrinking to fit.
- Encoding is asynchronous. The first image has no pixels until its encode lands. Later updates keep the displayed frame in place while the replacement is encoded. Its transmission is emitted before its native placeholders in one paint, so the host switches only after it has the pixels.
Captures
A capture holds images beside its cells rather than in them. TestBackend::capture_frame(), a UiSnapshot, and TerminalScreen::capture_frame() all fill CapturedFrame::images: one CapturedImage per image, with its RGBA pixels, the cells it is laid out over, and which of those cells still show it.
- What covers an image hides it. A UI capture records each image as the renderer draws it, then checks each of its cells after everything else has drawn. An overlay, a border, a toast, or a pane above takes the cells it covers, exactly as on the host.
CapturedImage::shows(x, y)answers per cell. - The cells get a half-block stand-in. Each visible cell holds
▀in the colors of its top and bottom halves, soplain_text()marks where an image is,to_ansi_text()shows a coarse version in any terminal, and cell assertions read colors straight out of the grid. A cell the image leaves fully transparent keeps what it held.CapturedImage::backgroundsrecords each cell's background from before the stand-in. - A PNG draws the pixels.
to_png()scales each image into its cells at the PNG's own cell size, keeping its aspect ratio from the top-left corner as a terminal does, and draws only the cells it still shows in. Transparent pixels show the recorded background, not the stand-in. - One image on its own.
CapturedImage::to_png()encodes justrgba, atwidthxheightwith its alpha, for a serializer that reports images beside the cells. A terminal capture has already cropped an image that runs past the viewport, so these are the pixels inside it. - A backdrop dims the pixels. Under an open modal backdrop a capture records the recolored pixels, so
rgba, the half-block stand-ins, andto_png()all show the image dimmed like the cells around it. So do the layers in the tree that reach pixels; the rest leave the image at full brightness where it still shows. - A capture does not wait for an encode. It never encodes for a host, so the first capture after an image arrives already has it.
TerminalScreen::capture_frame()crops each placement to the viewport the way the renderer does, and sizes its half blocks withcell_size().
Testing
Render through TestBackend and read capture_frame(): its images for pixels and visibility, its cells for the half-block colors. See tests/terminal_images_render.rs.