LLaVA is a novel end-to-end trained large multimodal model that combines a vision encoder and Vicuna
297
Sort by
TAG
Last pushed almost 2 years by mdelapenya
| Digest | OS/ARCH | Compressed size |
|---|---|---|
996f3eac202b | linux/amd64 | 5.93 GB |
2a782c00ee65 | linux/arm64 | 5.66 GB |