Rendered at 18:32:18 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
jerlendds 1 days ago [-]
VLM != LLM. Vision language models basically treat text tokens and image tokens the same. Post-training an LLM on images+text can improve its capabilities. Id recommend searching around the keyword VLM to find more resources on how multi-modal AI works.
transformers were first an image understanding technique, the text processing came later, it's all about the training data and gradient descent, and allegedly attention
laruss5 9 hours ago [-]
It's actually the other way round - the Transformer architecture was introduced for text (machine translation) in "Attention Is All You Need" (2017). Vision Transformers, which apply it to images, came three years later in 2020: https://arxiv.org/abs/2010.11929
verdverm 3 hours ago [-]
right, it was not text generation per-se (completion/contemporary understanding) that came first, vision was before that, translation before that
- https://huggingface.co/blog/vlms
- https://en.wikipedia.org/wiki/Multimodal_learning