All models

1 of 55 open-source models (filtered).

Sort:

Vision-language variant of Yi 34B. Image-text reasoning via an MLP adapter on a CLIP encoder. Useful for bilingual EN/中 multimodal workloads where the major Western vision-language models underperform on Chinese text in images.

Context: 4K
License: apache-2-0
VRAM Q4: 20.4 GB