Project updates and technical posts
Does the same model perform differently inside different agent harnesses? PawBench puts both model capability and harness quality into the same evaluation matrix so you can see how they jointly determine real-task outcomes.