How do you evaluate AI tools for UVM/DV work?

Hi everyone,

Sorry if this is a dumb question. I’m still pretty new to the verification side of AI tools and was hoping to learn from people here.

I’ve been trying Claude/GPT for things like UVM code generation and debugging, and sometimes it feels surprisingly good, while other times it confidently generates things that don’t even work. That made me realize I actually have no idea how people in industry evaluate these tools.

When you try a new model or AI assistant, how do you decide if it’s actually useful?

Do you have a few designs or testbenches that you always use?

Do you rely on benchmarks like CVDP or LBC-bench, or do those not matter much in practice?

Are there any blogs, papers, or people you’d recommend following for honest evaluations instead of vendor demos?

Thanks!