An MLX inference server for Apple Silicon with real-time telemetry
This project solves the 'black box' problem of running LLMs on macOS by exposing real tokens/sec and memory telemetry. It bundles Python and MLX into a native Swift menu bar app, removing the need for complex environment setup. It is ideal for Mac users who need to benchmark quantizations or monitor GPU memory ceilings accurately.
View on GitHub →canivel/mlx-dyno