trying to track every new release is a
total nightmare , but i found this guide that uses automationbench to see how models actually handle multi-step tasks. it tests things like gpt-5.6 sol and gemini 3.5 flash on
real workflows rather than just simple prompts.
it's way more useful than just reading marketing hype . anyone else found a better way to test these for automations?
found this here:
https://zapier.com/blog/ai-models-on-zapier