Pinned Loading
-
commvq-kv-cache-mlx
commvq-kv-cache-mlx PublicRoPE-commutative additive vector quantization of the KV cache in MLX, so quantized keys can be rotated without re-encoding.
Python
-
cppo-mlx
cppo-mlx PublicToy-scale MLX study of CPPO: pruning low-advantage completions in GRPO so each update needs far fewer sampled completions, plus dynamic completion allocation.
Python
-
delta-attention-lm
delta-attention-lm PublicA small hybrid LM in MLX mixing gated delta-rule attention, NoPE latent MoE and attention residuals over depth, after Kimi K3.
Python
-
-
kv-cache-fusion
kv-cache-fusion PublicCache-to-cache communication: a trainable fuser maps one small LM's KV cache into another's so it can answer with knowledge it never learned.
Python
-
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.