Skip to content
View original post on X: Lucas Beyer (bl16)· 7/100AI score7/100

Lucas Beyer proposes a small open multimodal image-text matching model

AISummary

Lucas Beyer says he could build a multimodal model that scores how well any freeform text fits a given image, with calibrated scores and sigmoid-based yes/no judgments. He floats raising roughly xxxM in funding, open-weighting the model, and keeping it under 1B parameters, perhaps around 400M.

Post on XView on X

Guys hear me out: Jev, but multimodal!

Imagine there was a model that you can give any image, and any set of freeform texts, and it would return you for each text, how well it fits the image.

It could also do a set of yes/no (Noul) by using sigmoid in the right places.

And the scores could be calibrated!

I think i should be able to do this, though i might need to raise xxxM. But as one of the inventors of vision encoders but also one of their biggest haters because i know better, i can pull this off.

And now imagine i would manage to open-weight it, wouldn't that be awesome?

And EVEN BETTER, imagine i could do it with fewer than 1B params?! Maybe like 400m or so?

I think I'm onto something here, wdyt??

Source: Lucas Beyer (bl16) · x.comPublished · added here