CRAB is an academic project that provides a benchmark for measuring how well multimodal language model agents perform when a task stretches across more than one environment. It comes from a team of researchers at several universities and labs, with a paper and code released publicly.
It is mainly useful for researchers and developers who build agents and want a consistent way to test and compare them. The project is free, and the code can be downloaded and extended.