The Spectrum Dispatch News

technology

MicroLLM Lab lets users benchmark tiny language models directly in their browser

An interactive tool lets people test seven small LLMs with objective measurements of speed and accuracy, generating shareable performance certificates.

MicroLLM Lab lets users benchmark tiny language models directly in their browser

MicroLLM Lab is an experimental platform that allows users to run and benchmark tiny language models directly in their web browser, according to the State of Utopia website.

MicroLLM Lab lets users benchmark tiny language models directly in their browser

The tool lets users select from seven small LLMs and execute performance tests on their own device. The benchmarking focuses on objective measurements rather than writing quality, using regex and exact token matching to evaluate model outputs. According to the platform, even a 135-million-parameter model is permitted to fail these tests as part of the measurement process.

Each test generates several key metrics. Runtime per test is measured in milliseconds, while accuracy is tracked as a pass rate on objective tests. The platform also calculates speed in tokens per second based on sustained decode performance, using either estimated values or previously recorded token-per-second rates from the user’s device.

Results are presented in real-time charts that display the latest benchmark suite data for each model. The interface shows accuracy as the latest suite pass rate, mean tokens per second for speed, and suite wall time in milliseconds.

One notable feature is the ability to generate and download a verifiable performance certificate. This certificate includes device hardware specifications, peak and sustained tokens-per-second measurements, and allows users to share their results on social media.

For developers interested in creating custom benchmarks, the platform provides a JavaScript editor that allows users to write their own benchmark tests. These checks are evaluated within the browser origin, and each check runs against the model’s decoded text output.

All performance data stays on the user’s machine rather than being collected centrally, according to the platform description. This approach keeps individual test results private while still allowing users to generate shareable certificates of their benchmark runs.

Key facts

  • Users can benchmark seven tiny LLMs directly in their browser
  • Tests measure runtime in milliseconds, accuracy via pass rates, and speed in tokens per second
  • Results generate downloadable verifiable performance certificates with hardware and speed data
  • Performance data stays on the user’s machine and is not collected centrally
  • Users can write custom benchmarks in JavaScript using the platform’s editor

Sources

← All posts