Latent

Applied research · started October 2026

Can a model edit a photo it barely looks at?

Latent is a project of 7xz. It is developing photo editing software built for language models: the program talks to the model in a form made for it, measurements and recipes instead of screenshots, to reach expert-level results at a small fraction of the usage.

Who is writing this? About me →

What it is for

Programs that communicate with language models in a way made for them, starting with photo editing that is precise and affordable.

Most tools reach a model through interfaces designed for people and wrapped afterwards. Latent starts from the other side: the model is the first user, so the program reports its state in compact numbers the model can act on, and the model works in a few well-defined operations instead of many low-level calls checked by screenshots.

The question

How close can a model get to expert retouching when it works from measurements and recipes instead of images, and what does each finished photo cost?

Status: early development. The evaluation is being set up before the method is built out.

The name. On film, a latent image forms the moment light hits it but stays invisible until it is developed. Here the model decides an edit it does not see, and the tool develops it. It pairs with Afterimage, which studies what a model keeps after it reads an input.

  1. Where it came from
  2. Why it is hard
  3. Approach
  4. How it will be judged
  5. What already exists
  6. Status and who

01 · Where it came from

I edit photos for albums with Claude Code driving GIMP through MCP: a dark, rich street-snap look, light retouching, sometimes light that falls like afternoon sun. It works, eventually. Two things went wrong along the way. Requests like "let the sunlight come in naturally" came out like a spotlight. And a handful of photos used up a large share of my usage.

I measured one of those sessions, two days of editing in a single conversation:

The cost did not come from any single edit. It came from a conversation that kept growing, full of screenshots, and was read again on every response. Caching made each read cheap, but not cheap enough at that volume.

02 · Why it is hard

A language model sees an image as a few hundred to a couple of thousand tokens made from small patches. That is good for what is in the picture and where. It is weak at exact values: how bright, how warm, how much of the highlights is clipped. So the model adjusts by trial and error and checks each step by looking again, and every look costs tokens and stays in the conversation.

Some of the work cannot be reduced to numbers. Whether light looks natural, whether a face still looks like the person, whether a retouch shows: those need judgment, and for now that means an image.

03 · Approach

04 · How it will be judged

My taste is one person's taste, so it cannot be the measure. The method has to be judged on problems with answers:

The target is to match the screenshot approach on quality at a small fraction of its cost.

05 · What already exists

06 · Status and who

Latent is a project of 7xz, a one-person studio in Korea, in early development: a measured baseline, design principles and a plan for evaluation. I build it with Claude Code.

Photo editing is the first case, not the whole plan. If the approach holds up there, the next step is to carry the same interface to other work where models now act through screenshots and wrapped tools: video color and editing first, then other visual programs. The aim is a common way for programs to talk to language models, proven one domain at a time.

It shares a question with Afterimage. Afterimage asks how a model takes in an input on the inside; Latent asks how far quality goes when the input is given as numbers and structure instead of pixels.

Contact: me@7xz.dev

Sources

  1. JarvisArt: an MLLM photo retouching agent for Lightroom (NeurIPS 2025)
  2. MIT-Adobe FiveK dataset
  3. Imagen AI developer API
  4. Claude vision: how images are turned into tokens
  5. Claude API pricing, including prompt caching
  6. Code execution with MCP (overview)
  7. mcp2cli: MCP servers as CLI tools
  8. Afterimage, the companion research project