DeepSeek Quietly Rolls Out New Vision-Language Model
DeepSeek has published documentation for a new model called v4-flash-vision-exp, signaling the Chinese AI lab's continued push into multimodal territory. The model appears to extend DeepSeek's fast 'flash' tier with the ability to process and reason about images alongside text, based on the API guide now live on its developer docs site.
Details remain sparse since this looks like an experimental release rather than a full production launch. There's no accompanying blog post with benchmarks, pricing, or a detailed changelog, just the technical documentation for developers who want to start testing vision inputs through the API.
DeepSeek has built a reputation for shipping capable models at aggressive price points, undercutting Western labs like OpenAI and Anthropic on cost while staying competitive on benchmarks. Adding vision to a lightweight 'flash' variant suggests they're targeting developers who want cheap, fast multimodal inference rather than frontier-level image reasoning.