StarVector is a foundation model for SVG generation that transforms vectorization into a code generation task. Using a vision-language modeling architecture, StarVector processes both visual and textual inputs to produce high-quality SVG code with remarkable precision.
A cross-platform video structuring (video analysis) framework based on CV models & mLLM.
✨✨Woodpecker: Hallucination Correction for Multimodal Large Language Models
Cambrian-S: Towards Spatial Supersensing in Video
🔥🔥🔥 A curated list of papers on LLMs-based multimodal generation (image, video, 3D and audio).