Applied research · started October 2026
Can a model edit a photo it barely looks at?
Latent is a project of 7xz. It is developing photo editing software built for language models: the program talks to the model in a form made for it, measurements and recipes instead of screenshots, to reach expert-level results at a small fraction of the usage.
Who is writing this? About me →What it is for
Programs that communicate with language models in a way made for them, starting with photo editing that is precise and affordable.
Most tools reach a model through interfaces designed for people and wrapped afterwards. Latent starts from the other side: the model is the first user, so the program reports its state in compact numbers the model can act on, and the model works in a few well-defined operations instead of many low-level calls checked by screenshots.
The question
How close can a model get to expert retouching when it works from measurements and recipes instead of images, and what does each finished photo cost?
Status: early development. The evaluation is being set up before the method is built out.
The name. On film, a latent image forms the moment light hits it but stays invisible until it is developed. Here the model decides an edit it does not see, and the tool develops it. It pairs with Afterimage, which studies what a model keeps after it reads an input.
01 · Where it came from
I edit photos for albums with Claude Code driving GIMP through MCP: a dark, rich street-snap look, light retouching, sometimes light that falls like afternoon sun. It works, eventually. Two things went wrong along the way. Requests like "let the sunlight come in naturally" came out like a spotlight. And a handful of photos used up a large share of my usage.
I measured one of those sessions, two days of editing in a single conversation:
- 1,808 model responses and over 500 GIMP calls (curves, color balance, hue and saturation, vignette, blur, and many raw API calls).
- 252 images returned into the conversation, mostly screenshots for checking the result.
- About 755 million input tokens read back from cache, plus about 2 million output tokens.
- The conversation read on each response grew from about 116,000 tokens on average in the first hour to about 695,000 on the last day.
The cost did not come from any single edit. It came from a conversation that kept growing, full of screenshots, and was read again on every response. Caching made each read cheap, but not cheap enough at that volume.
02 · Why it is hard
A language model sees an image as a few hundred to a couple of thousand tokens made from small patches. That is good for what is in the picture and where. It is weak at exact values: how bright, how warm, how much of the highlights is clipped. So the model adjusts by trial and error and checks each step by looking again, and every look costs tokens and stays in the conversation.
Some of the work cannot be reduced to numbers. Whether light looks natural, whether a face still looks like the person, whether a retouch shows: those need judgment, and for now that means an image.
03 · Approach
- An interface made for the model. The program is designed around how a language model reads and acts, not adapted from a screen made for people.
- Keep the state out of the conversation. The edit lives in a project file. The model does not re-read the whole history on each step.
- Measure instead of looking. After each step the tool reports numbers: brightness, contrast, black and white points, clipping, saturation, color temperature, sharpness, and the same inside and outside a selection. About a hundred tokens instead of a screenshot.
- Recipes instead of improvising. A look is a handful of parameters with a known effect. The model picks a recipe and adjusts a few values rather than assembling dozens of low-level calls.
- Look only when it matters. A small crop for judgments numbers cannot make, and one final check.
- A lighter interface. A command-line tool with a short guide instead of a large MCP tool list loaded into every conversation.
04 · How it will be judged
My taste is one person's taste, so it cannot be the measure. The method has to be judged on problems with answers:
- Reproducing expert edits. How close it gets to retouching done by professionals. A candidate is the MIT-Adobe FiveK dataset, 5,000 raw photos each edited by five experts; its research license still has to be checked for this use.
- Following instructions. How well it carries out requests like "warmer, light from the left, keep the highlights", judged blind.
- Cost per photo. Tokens per finished edit, against a model that checks with screenshots.
The target is to match the screenshot approach on quality at a small fraction of its cost.
05 · What already exists
- JarvisArt trains a multimodal model to drive more than 200 Lightroom tools. It looks at the image; this project tries to look as little as possible.
- Imagen AI offers a commercial API that learns a photographer's style and applies it to whole shoots.
- Code execution and CLI wrappers for MCP already cut the token cost of tool definitions and large results. They make each call cheaper but do not tell the model whether the edit worked.
- Generative image models can edit from a sentence, but they redraw the picture rather than adjust it, which is not what precise retouching needs.
06 · Status and who
Latent is a project of 7xz, a one-person studio in Korea, in early development: a measured baseline, design principles and a plan for evaluation. I build it with Claude Code.
Photo editing is the first case, not the whole plan. If the approach holds up there, the next step is to carry the same interface to other work where models now act through screenshots and wrapped tools: video color and editing first, then other visual programs. The aim is a common way for programs to talk to language models, proven one domain at a time.
It shares a question with Afterimage. Afterimage asks how a model takes in an input on the inside; Latent asks how far quality goes when the input is given as numbers and structure instead of pixels.
Contact: me@7xz.dev
Sources
- JarvisArt: an MLLM photo retouching agent for Lightroom (NeurIPS 2025)
- MIT-Adobe FiveK dataset
- Imagen AI developer API
- Claude vision: how images are turned into tokens
- Claude API pricing, including prompt caching
- Code execution with MCP (overview)
- mcp2cli: MCP servers as CLI tools
- Afterimage, the companion research project