Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I was also under the impression modern AI agents have moved on from just OCR'ing screenshots to leveraging native vision model capabilities.


They do. They all use ViTs and have for quite a while.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: