cryptopoly / ChaosEngineAI Sponsor Star 24 Code Issues Pull requests Local AI workstation — discover, run, chat, benchmark, and generate images from open-weight models. DFlash/DDTree speculative decoding, TurboQuant & TriAttention cache compression strategies, MLX + llama.cpp + vLLM + MTPLX backends. desktop-app python machine-learning typescript ai image-generation mlx tauri huggingface apple-silicon openai-api cache-compression llm stable-diffusion llama-cpp vllm local-ai gguf speculative-decoding dflash Updated Aug 1, 2026 Python
Labeeb2339 / recurquant Star 0 Code Issues Pull requests Experiments with mixed INT4/INT8 storage for Qwen3.5 recurrent state. research transformers pytorch quantization mixed-precision linear-attention cache-compression llm-inference qwen3 qwen35 gated-deltanet recurrent-state Updated Aug 14, 2026 Python