TokenLight steers photo relighting with five physical attributes — and holds up where it was tested
October 5, 2026 3 min read
TokenLight reformulates image relighting as a conditional generation problem. The model takes a photograph plus a small set of attribute tokens — encodings of light intensity, color, ambient illumination, diffuse level, and light position — and directly generates a relit version of that photograph. The control surface behaves like the light sliders in a 3D application, applied to flat images.
How does the method work?
Each lighting attribute is encoded as its own token, and because every attribute is separate and individually adjustable, edits compose: you can change a light's color without moving it, or dim the ambient while leaving a fixture untouched.
Training rests on two data sources. The bulk is synthetic — paired lighting examples rendered under systematically varied light, which carry known values for each attribute. A smaller set of real captures, taken indoors with visible fixtures toggled on and off, closes the gap between rendered and photographic appearance. Three practical operations fall out of the same representation: placing virtual lights anywhere in space, editing global ambient and diffuse properties, and switching individual in-scene fixtures on or off.
What does the paper establish?
Within its tested scope, the token controls produce plausible relighting on still images — visible fixtures and placed virtual lights — and indoor scenes hold together better than outdoor ones. The paper also compares its outputs against other relighting methods. A caveat belongs to that comparison: the compared methods are closed-source, so their outputs were supplied by their respective authors rather than reproduced independently, and the LightLab examples came from its own publication. Because the relative results rest on the authors' runs, the ranking is theirs to stand behind; the checked scope does not reproduce it.
What does it not establish?
- Non-realtime. The paper reports that large-model inference makes immediate interactive feedback challenging, so expect non-realtime generation rather than assuming live preview is available.
- Seed stability. Outputs vary across random seeds, so identical inputs and settings yield slightly different relights.
- Outdoor generalization. Indoor scenes generalize better than outdoor ones, and the real-capture anchor of training is indoor fixture photography.
- Time and motion. Video relighting with moving cameras and objects is named as an open problem; because the method operates on single photographs, nothing in the checked scope claims consistency across views or frames.
- Transfer. Nothing in the checked scope addresses extending the token scheme to other architectures, tasks, or media.
Does this matter for your task?
The applicability condition is concrete: if your pipeline relights still images offline — interiors with identifiable fixtures — and you can absorb non-realtime generation plus seed-to-seed variation (generate several candidates, keep the best), the attribute-token interface is a promising control surface. Whether it is meaningfully cleaner than text prompts or environment maps is my interpretation, not a measured advantage; the checked scope does not rank interfaces head to head. If you need drag-and-drop live lighting, strong outdoor performance, or anything temporal, the paper's own limitations say the method is not there yet — and nothing in it licenses assuming otherwise.
Sources
- TokenLight: Precise Lighting Control in Images using Attribute Tokens — arXiv (author-submitted research)