New Developments

Shangtech unveils open-source multimodal model with native 4k image output

SenseNovaSource: IT之家 (ITHome)21/08/2026, 08:09
ShangTech has announced the open-sourcing of its lightweight, native multimodal large model, SenseNova U1.5 Lite. The model supports a context length of 3-4K and can handle multiple constraints such as subject, quantity, spatial relationships, text, layout, and style, enhancing the stability of complex visual tasks. It features improved visual generation with better composition, color, texture, lighting, realism, and fine details, reducing instances where local accuracy does not match overall completion. The model also offers more reliable native image editing capabilities, including enhanced subject identity, spatial structure, layout relationships, and preservation of non-edited areas, along with improved local modifications, element replacement, text refinement, and multi-reference image editing. SenseNova U1.5 Lite strengthens its ability to handle complex layouts, including Chinese and English text, posters, infographics, brand visuals, and multi-text formatting, moving from content generation to organizing complete visual expressions. It supports precise visual control through bounding boxes, visual markers, and single or multiple reference images. The model also provides native 4K high-resolution output, balancing overall composition with fine textures, small text, and light refraction. As an 8B parameter model, SenseNova U1.5 Lite outperforms other models of similar scale in instruction following and image editing consistency, achieving performance comparable to large commercial models in text rendering and complex layouts. The model is available on GitHub, Hugging Face, and ModelScope.
Shangtech unveils open-source multimodal model with native 4k image output — lupAI