{"name":"project-maya","repo":"mw00/project-maya","url":"https://midorreal.com/project/mw00-project-maya","repository":"https://github.com/mw00/project-maya","description":"Project Maya - GLM-5.3-Flash (321B MoE) on your own NVIDIA GPU(s): engine, server, dashboard, quant tools. Built on Strata.","facts":{"license":{"value":"MIT","source":"https://github.com/mw00/project-maya","checked":"9 Oct 2026"},"language":{"value":"C++","source":"https://github.com/mw00/project-maya","checked":"9 Oct 2026"},"last_release":{"value":"v1.0.16, 8 Oct 2026","source":"https://github.com/mw00/project-maya/releases","checked":"9 Oct 2026"},"releases":{"value":"17","source":"https://github.com/mw00/project-maya/releases","checked":"9 Oct 2026"},"contributors":{"value":"6","source":"https://github.com/mw00/project-maya/graphs/contributors","checked":"9 Oct 2026"},"archived":{"value":"no","source":"https://github.com/mw00/project-maya","checked":"9 Oct 2026"}},"stars":151.0,"writeup":{"what_it_is":"Project Maya runs GLM-5.3-Flash, a 321-billion-parameter mixture-of-experts AI model, on consumer NVIDIA GPUs by distributing model layers across GPU VRAM, system RAM, and NVMe SSD. It provides a chat interface, image support, and OpenAI/Anthropic-compatible APIs.","audience":"Users with NVIDIA GPUs (V100 or newer) wanting to run large language models locally.","claims":[{"kind":"specific","claim":"Runs on a single GPU or up to 16 GPUs that share the model","excerpt":"one GPU, or up to 16 that share the model","status":"not_checked"},{"kind":"specific","claim":"Achieves up to 40 tokens/s decode speed on 2x Tesla V100 32GB GPUs","excerpt":"**up to 40 tokens/s**","status":"not_checked"},{"kind":"specific","claim":"Requires approximately 100 GB free on a fast NVMe SSD","excerpt":"~100 GB free on a fast NVMe SSD","status":"not_checked"},{"kind":"specific","claim":"Supports context lengths up to 1 million tokens","excerpt":"a context of up to 1 M tokens","status":"not_checked"}],"alternatives":[{"name":"omlx","url":"https://midorreal.com/project/jundot-omlx"},{"name":"ODS","url":"https://midorreal.com/project/osmantic-ods"}],"written_by":"AI, from the project's README","date":"2026-10-09"},"why_now":null,"owner_supplied":null,"score":{"verdict":"real","hype":38,"reality":63,"gap":-25,"momentum":null,"confidence":60,"version":"score-2.0","rule":"verdict-2.0","calculated_at":"2026-10-09T06:00:00+00:00","verdict_source":"rule"},"early_read":null,"changes":[{"date":"2026-10-08T21:39:25+00:00","text":"Latest release v1.0.16","source":"https://github.com/mw00/project-maya/releases"},{"date":"2026-10-07T20:38:04+00:00","text":"First release","source":"https://github.com/mw00/project-maya/releases"},{"date":"2026-10-06T22:55:13+00:00","text":"Repository created","source":"https://github.com/mw00/project-maya"}],"cite":"project-maya is a GitHub project at https://github.com/mw00/project-maya written mainly in C++, described there as \"Project Maya - GLM-5.3-Flash (321B MoE) on your own NVIDIA GPU(s): engine, server, dashboard, quant tools. Built on Strata\".","attribution":"Data from Mid or Real, https://midorreal.com/project/mw00-project-maya, checked 9 October 2026. Attribution with a link is required."}