The AI Competency Perception Ladder
A fascinating part of R&D has been the progression of how we anthropomorphically judge AI models, not just LLMs - synthetic voice, language translation, GPT, image gen, video gen, coding.
Cultural critique always goes like:
- there are glimmers of useful, robotic functionality (but it’ll never do XYZ)
- it’s hilariously bad (which means it’s passed some bio human benchmark and now triggers our “good v bad” radar)
- it’s robotically bad (this is where it starts to draw tons of critique bc humans almost register it as an inferior being and sense workforce potential)
- it’s error-prone (we can see and measure the mistakes like we regularly do with humans)
- it feels bad (passably human, but with cultural mistakes. Critiqued based on taste and emotional preference, or “je ne sais quoi”)
- with experienced evaluation, you can see the flaws (experts are paying attention, and biased detractors collect mistakes)
- really good, but which company’s does it best?
It’s funny looking back at my notes on robot voices, or image gen, waiting for something useable, and now it’s just a vendor choice.