We were a little impatient. Qwen 3.8 27B was released on August 14, 2026. Four days later, on August 18, we published our first set of GGUFs. We called them ShapeLearn-Lite for a reason: they were produced using a much smaller optimization budget, fewer checks, and much less waiting. Now the full ShapeLearn models are done, and we have benchmarked them alongside the original Lite set and competing quants. The good news: ShapeLearn-Lite held up pretty well. We will come back to that later in “ShapeLearn-Lite, in retrospect”. The better news: the full ShapeLearn models are even better.
Quick start with llama.cpp The MTP draft head is bundled in every GGUF. DFlash2 uses a separate 1.1 GB draft model. Both commands use GPU-5 . Swap the tag for any other model in the release. MTP Embedded draft. Works with image inputs. Copy llama-server \ -hf byteshape/Qwen3.8-27B-GGUF:Qwen3.8-27B-IQ4_XS-3.84bpw \ --mmproj-auto \ --spec-type draft-mtp --spec-draft-n-max 3 DFlash2 External draft. Fastest option, text only. Copy llama-server \ -hf byteshape/Qwen3.8-27B-GGUF:Qwen3.8-27B-IQ4_XS-3.84bpw \ -hfd incoai/Qwen3.8-27B-DFlash2-GGUF:Q4_K_M \ --spec-type draft-dflash --spec-draft-n-max 7 \ --no-mmproj DFlash2 needs llama.cpp b10658 or newer. Ready-to-run commands for every model, with the recommended sampling settings, are in the run tool and on the model card.
TL;DR Full ShapeLearn moves the measured quality-speed frontier beyond Lite. All five models in the new release sit on the frontier in each of our six GPU comparisons.
GPU-5 is our default recommendation wherever it fits, reaching 99.63% of BF16’s aggregate benchmark score. If it does not fit with the context you need, GPU-4 is still very competitive: it reaches 98.72% of BF16 at a much smaller size (11.0 GB instead of 13.1 GB), and it is faster.
is our default recommendation wherever it fits, reaching 99.63% of BF16’s aggregate benchmark score. If it does not fit with the context you need, is still very competitive: it reaches 98.72% of BF16 at a much smaller size (11.0 GB instead of 13.1 GB), and it is faster. ShapeLearn-Lite also performed better than its KLD ranking suggested: three of its six models sit on the frontier in the Lite-versus-Unsloth Dynamic v3 comparison.
Speculative Decoding with MTP or DFlash2 increases throughput across every ShapeLearn model and GPU tested. DFlash2 is usually faster but requires more memory and does not support image inputs with llama.cpp. Choose DFlash2 for maximum text-only throughput when memory allows, and MTP when VRAM or multimodal support matters more.
Full ShapeLearn moves the frontier We are releasing the full ShapeLearn run for Qwen 3.8 27B. Within this release, larger models yield higher aggregate scores, while smaller models deliver higher throughput. That ordering holds across all six GPUs tested. Because this is a dense model and memory transfers are the bottleneck, lower BPW translates more directly into higher TPS than it does for MoEs. The per-GPU comparisons also include AtomicChat, Bartowski, ISTA-DASLab, and Unsloth Dynamic v3. Bartowski’s latest models were released after our testing and are not included. Full ShapeLearn is labelled ByteShape in the figures. All five ShapeLearn models remain on the measured frontier, with GPU-5 achieving the highest aggregate score among the plotted quants. Other teams also contribute competitive points. Notably, ISTA-DASLab’s excellent model (the yellow “d” on the graph below) also sits on the frontier. By “frontier,” we mean that no other plotted model is both faster and more accurate. 96 GB: RTX Pro 6000 RTX PRO 6000, the GPU with the most memory, lets us show the full range of models tested. RTX Pro 6000: tokens per second vs quality (full ShapeLearn and competing quants) Tap Show Legend below for model details. RTX Pro 6000: tokens per second vs quality (full ShapeLearn and competing quants) Hover over the bubbles, or click Show Legend below, for model details. Show Legend # Model Acc TPS BPW ByteShape GPU-1 IQ2_XXS-2.56bpw 0.9304 116.11 2.56 GPU-2 IQ3_XXS-2.88bpw 0.9656 108.01 2.88 GPU-3 IQ3_XS-3.01bpw 0.9726 105.44 3.01 GPU-4 IQ3_S-3.23bpw 0.9872 101.11 3.23 GPU-5 IQ4_XS-3.84bpw 0.9963 90.42 3.84 Unsloth A UD-IQ2_S 0.8633 114.81 2.49 B UD-Q2_K_XL 0.9572 106.95 2.82 C UD-IQ3_XXS 0.9359 100.51 3.14 D UD-IQ3_S 0.9555 95.53 3.47 E UD-Q3_K_XL 0.9760 90.88 3.80 F UD-IQ4_XS 0.9920 86.73 4.13 G UD-Q4_K_S 0.9877 82.21 4.46 H UD-Q4_K_M 0.9703 78.58 4.79 I UD-Q4_K_XL 0.9871 74.75 5.12 J UD-Q5_K_S 0.9878 71.00 5.44 K UD-Q5_K_M 0.9897 67.71 5.77 L UD-Q5_K_XL 0.9905 65.52 6.10 ISTA-DASLab a GSQ-RCO-IQ2_XS 0.8647 112.77 2.50 b GSQ-RCO-IQ2_S 0.9364 108.22 2.75 c GSQ-RCO-IQ3_XXS 0.9438 103.39 3.00 d GSQ-RCO-IQ3_S 0.9943 94.77 3.50 Bartowski a IQ2_XXS 0.7986 112.21 2.72 b IQ2_S 0.9296 106.33 2.99 c Q2_K 0.9616 96.08 3.45 d IQ3_XXS 0.9594 92.43 3.68 e IQ3_XS 0.9582 88.36 3.89 f IQ3_M 0.9667 86.17 4.06 AtomicChat a AD-IQ2_XXS 0.8385 116.28 2.58 b AD-IQ2_XS 0.9296 108.67 2.85 c AD-IQ2_S 0.9061 100.28 3.22 d AD-IQ3_XXS 0.9644 95.29 3.50 e AD-IQ3_S 0.9730 88.52 4.04 GPU-5 is our default wherever you can fit it: it reaches 90.4 tok/s at 99.63% of the BF16 baseline. 32 GB: RTX 5090 The RTX 5090 tells a similar story, leading to the same recommendations. RTX 5090: tokens per second vs quality (full ShapeLearn and competing quants) Tap Show Legend below for model details. RTX 5090: tokens per second vs quality (full ShapeLearn and competing quants) Hover over the bubbles, or click Show Legend below, for model details. Show Legend # Model Acc TPS BPW ByteShape GPU-1 IQ2_XXS-2.56bpw 0.9304 119.13 2.56 GPU-2 IQ3_XXS-2.88bpw 0.9656 110.78 2.88 GPU-3 IQ3_XS-3.01bpw 0.9726 108.08 3.01 GPU-4 IQ3_S-3.23bpw 0.9872 103.57 3.23 GPU-5 IQ4_XS-3.84bpw 0.9963 93.66 3.84 Unsloth A UD-IQ2_S 0.8633 115.34 2.49 B UD-Q2_K_XL 0.9572 108.12 2.82 C UD-IQ3_XXS 0.9359 102.87 3.14 D UD-IQ3_S 0.9555 97.87 3.47 E UD-Q3_K_XL 0.9760 93.46 3.80 F UD-IQ4_XS 0.9920 89.69 4.13 G UD-Q4_K_S 0.9877 85.28 4.46 H UD-Q4_K_M 0.9703 81.63 4.79 I UD-Q4_K_XL 0.9871 77.51 5.12 J UD-Q5_K_S 0.9878 73.60 5.44 K UD-Q5_K_M 0.9897 69.87 5.77 L UD-Q5_K_XL 0.9905 67.59 6.10 ISTA-DASLab a GSQ-RCO-IQ2_XS 0.8647 114.11 2.50 b GSQ-RCO-IQ2_S 0.9364 110.19 2.75 c GSQ-RCO-IQ3_XXS 0.9438 105.67 3.00 d GSQ-RCO-IQ3_S 0.9943 95.49 3.50 Bartowski a IQ2_XXS 0.7986 115.77 2.72 b IQ2_S 0.9296 109.39 2.99 c Q2_K 0.9616 99.29 3.45 d IQ3_XXS 0.9594 95.83 3.68 e IQ3_XS 0.9582 91.14 3.89 f IQ3_M 0.9667 89.08 4.06 AtomicChat a AD-IQ2_XXS 0.8385 119.53 2.58 b AD-IQ2_XS 0.9296 112.39 2.85 c AD-IQ2_S 0.9061 103.46 3.22 d AD-IQ3_XXS 0.9644 98.22 3.50 e AD-IQ3_S 0.9730 91.72 4.04 Once again GPU-5 is our default choice, reaching 93.7 tok/s. Choose GPU-4 for slightly more context length or slightly better TPS. 24 GB: RTX 4090 and RTX 3090 Both 24 GB cards fit all five ShapeLearn models. We plot them separately because their throughput differs, but the ordering is the same on both. RTX 4090 The RTX 4090 keeps the same pattern: GPU-5 is the default, reaching 59.2 tok/s. RTX 4090: tokens per second vs quality (full ShapeLearn and competing quants) Tap Show Legend below for model details. RTX 4090: tokens per second vs quality (full ShapeLearn and competing quants) Hover over the bubbles, or click Show Legend below, for model details. Show Legend # Model Acc TPS BPW ByteShape GPU-1 IQ2_XXS-2.56bpw 0.9304 78.67 2.56 GPU-2 IQ3_XXS-2.88bpw 0.9656 72.89 2.88 GPU-3 IQ3_XS-3.01bpw 0.9726 71.12 3.01 GPU-4 IQ3_S-3.23bpw 0.9872 67.77 3.23 GPU-5 IQ4_XS-3.84bpw 0.9963 59.16 3.84 Unsloth A UD-IQ2_S 0.8633 78.23 2.49 B UD-Q2_K_XL 0.9572 72.79 2.82 C UD-IQ3_XXS 0.9359 67.24 3.14 D UD-IQ3_S 0.9555 63.02 3.47 E UD-Q3_K_XL 0.9760 59.13 3.80 F UD-IQ4_XS 0.9920 55.62 4.13 G UD-Q4_K_S 0.9877 52.44 4.46 H UD-Q4_K_M 0.9703 49.77 4.79 I UD-Q4_K_XL 0.9871 47.08 5.12 J UD-Q5_K_S 0.9878 44.80 5.44 K UD-Q5_K_M 0.9897 42.45 5.77 L UD-Q5_K_XL 0.9905 40.84 6.10 ISTA-DASLab a GSQ-RCO-IQ2_XS 0.8647 77.80 2.50 b GSQ-RCO-IQ2_S 0.9364 73.34 2.75 c GSQ-RCO-IQ3_XXS 0.9438 69.26 3.00 d GSQ-RCO-IQ3_S 0.9943 62.41 3.50 Bartowski a IQ2_XXS 0.7986 75.28 2.72 b IQ2_S 0.9296 71.27 2.99 c Q2_K 0.9616 63.74 3.45 d IQ3_XXS 0.9594 61.03 3.68 e IQ3_XS 0.9582 58.20 3.89 f IQ3_M 0.9667 56.16 4.06 AtomicChat a AD-IQ2_XXS 0.8385 79.51 2.58 b AD-IQ2_XS 0.9296 74.19 2.85 c AD-IQ2_S 0.9061 67.40 3.22 d AD-IQ3_XXS 0.9644 63.60 3.50 e AD-IQ3_S 0.9730 57.38 4.04 RTX 3090 Older, but still fast in these measurements. RTX 3090: tokens per second vs quality (full ShapeLearn and competing quants) Tap Show Legend below for model details. RTX 3090: tokens per second vs quality (full ShapeLearn and competing quants) Hover over the bubbles, or click Show Legend below, for model details. Show Legend # Model Acc TPS BPW ByteShape GPU-1 IQ2_XXS-2.56bpw 0.9304 53.03 2.56 GPU-2 IQ3_XXS-2.88bpw 0.9656 51.20 2.88 GPU-3 IQ3_XS-3.01bpw 0.9726 50.75 3.01 GPU-4 IQ3_S-3.23bpw 0.9872 49.49 3.23 GPU-5 IQ4_XS-3.84bpw 0.9963 45.79 3.84 Unsloth A UD-IQ2_S 0.8633 52.82 2.49 B UD-Q2_K_XL 0.9572 50.37 2.82 C UD-IQ3_XXS 0.9359 48.16 3.14 D UD-IQ3_S 0.9555 46.39 3.47 E UD-Q3_K_XL 0.9760 46.28 3.80 F UD-IQ4_XS 0.9920 46.12 4.13 G UD-Q4_K_S 0.9877 44.49 4.46 H UD-Q4_K_M 0.9703 43.30 4.79 I UD-Q4_K_XL 0.9871 41.37 5.12 J UD-Q5_K_S 0.9878 39.37 5.44 K UD-Q5_K_M 0.9897 37.38 5.77 L UD-Q5_K_XL 0.9905 36.16 6.10 ISTA-DASLab a GSQ-RCO-IQ2_XS 0.8647 51.56 2.50 b GSQ-RCO-IQ2_S 0.9364 50.07 2.75 c GSQ-RCO-IQ3_XXS 0.9438 48.66 3.00 d GSQ-RCO-IQ3_S 0.9943 47.19 3.50 Bartowski a IQ2_XXS 0.7986 53.63 2.72 b IQ2_S 0.9296 51.21 2.99 c Q2_K 0.9616 45.34 3.45 d IQ3_XXS 0.9594 47.46 3.68 e IQ3_XS 0.9582 44.68 3.89 f IQ3_M 0.9667 43.42 4.06 AtomicChat a AD-IQ2_XXS 0.8385 53.95 2.58 b AD-IQ2_XS 0.9296 51.91 2.85 c AD-IQ2_S 0.9061 48.24 3.22 d AD-IQ3_XXS 0.9644 47.10 3.50 e AD-IQ3_S 0.9730 47.22 4.04 GPU-4 reaches 49.5 tok/s, compared with 45.8 tok/s for GPU-5 . Moving to the larger model costs about 7.5% in throughput, while the aggregate score rises from 98.72% to 99.63% of BF16. That makes GPU-5 the default here as well. 16 GB: RTX 4080 and RTX 5060 Ti With a tighter VRAM budget, the pragmatic choice is to leave room for the context you need, not just the model weights. These plots contain fewer competing configurations, but all five ShapeLearn models are represented. RTX 4080 On the RTX 4080, GPU-4 reaches 52.4 tok/s, while GPU-5 reaches 45.7 tok/s. RTX 4080: tokens per second vs quality (full ShapeLearn and competing quants) Tap Show Legend below for model details. RTX 4080: tokens per second vs quality (full ShapeLearn and competing quants) Hover over the bubbles, or click Show Legend below, for model details. Show Legend # Model Acc TPS BPW ByteShape GPU-1 IQ2_XXS-2.56bpw 0.9304 62.01 2.56 GPU-2 IQ3_XXS-2.88bpw 0.9656 56.84 2.88 GPU-3 IQ3_XS-3.01bpw 0.9726 55.14 3.01 GPU-4 IQ3_S-3.23bpw 0.9872 52.43 3.23 GPU-5 IQ4_XS-3.84bpw 0.9963 45.74 3.84 Unsloth A UD-IQ2_S 0.8633 61.70 2.49 B UD-Q2_K_XL 0.9572 56.70 2.82 C UD-IQ3_XXS 0.9359 52.32 3.14 D UD-IQ3_S 0.9555 49.06 3.47 ISTA-DASLab a GSQ-RCO-IQ2_XS 0.8647 60.47 2.50 b GSQ-RCO-IQ2_S 0.9364 57.37 2.75 c GSQ-RCO-IQ3_XXS 0.9438 54.19 3.00 d GSQ-RCO-IQ3_S 0.9943 48.42 3.50 Bartowski a IQ2_XXS 0.7986 58.69 2.72 b IQ2_S 0.9296 55.39 2.99 c Q2_K 0.9616 48.91 3.45 d IQ3_XXS 0.9594 46.99 3.68 e IQ3_XS 0.9582 44.71 3.89 AtomicChat a AD-IQ2_XXS 0.8385 62.55 2.58 b AD-IQ2_XS 0.9296 57.82 2.85 c AD-IQ2_S 0.9061 52.43 3.22 d AD-IQ3_XXS 0.9644 49.15 3.50 RTX 5060 Ti On the RTX 5060 Ti, the corresponding figures are 33.1 tok/s and 29.1 tok/s. RTX 5060 Ti: tokens per second vs quality (full ShapeLearn and competing quants) Tap Show Legend below for model details. RTX 5060 Ti: tokens per second vs quality (full ShapeLearn and competing quants) Hover over the bubbles, or click Show Legend below, for model details. Show Legend # Model Acc TPS BPW ByteShape GPU-1 IQ2_XXS-2.56bpw 0.9304 38.05 2.56 GPU-2 IQ3_XXS-2.88bpw 0.9656 35.41 2.88 GPU-3 IQ3_XS-3.01bpw 0.9726 34.50 3.01 GPU-4 IQ3_S-3.23bpw 0.9872 33.05 3.23 GPU-5 IQ4_XS-3.84bpw 0.9963 29.15 3.84 Unsloth A UD-IQ2_S 0.8633 38.02 2.49 B UD-Q2_K_XL 0.9572 35.42 2.82 C UD-IQ3_XXS 0.9359 32.80 3.14 D UD-IQ3_S 0.9555 30.97 3.47 ISTA-DASLab a GSQ-RCO-IQ2_XS 0.8647 37.62 2.50 b GSQ-RCO-IQ2_S 0.9364 35.80 2.75 c GSQ-RCO-IQ3_XXS 0.9438 33.95 3.00 d GSQ-RCO-IQ3_S 0.9943 30.72 3.50 Bartowski a IQ2_XXS 0.7986 36.33 2.72 b IQ2_S 0.9296 34.76 2.99 c Q2_K 0.9616 30.79 3.45 d IQ3_XXS 0.9594 29.69 3.68 e IQ3_XS 0.9582 28.24 3.89 AtomicChat a AD-IQ2_XXS 0.8385 38.42 2.58 b AD-IQ2_XS 0.9296 36.05 2.85 c AD-IQ2_S 0.9061 32.61 3.22 d AD-IQ3_XXS 0.9644 31.01 3.50 GPU-5 remains the default on both cards when the model, KV cache, and runtime buffers fit within your memory budget. When they do not, GPU-4 is still very competitive: almost 99% of BF16 at a much smaller size, and faster. A model appearing in these measurements does not establish that every context length or serving configuration will fit.
ShapeLearn-Lite, in retrospect ShapeLearn-Lite uses a smaller optimization budget than full ShapeLearn. It let us get Qwen 3.8 27B onto 12 GB to 24 GB GPUs within a few days. We released after targeted sanity checks and started the full evaluation afterwards. The full ShapeLearn models were ready before the benchmarking was finished. Evaluating both sets, along with the competing models, is what took most of the time. Then Unsloth released its Dynamic v3 models. At similar sizes, several had lower KLD than Lite in our measurements. On KLD alone, Lite looked less competitive. KLD looked decisive KLD measures divergence between a quantized model’s predicted token distributions and the BF16 reference under a particular evaluation setup. It is useful for diagnosing substantial changes, but lower divergence does not automatically mean better task performance. We measure KLD on a dataset of about 5 million tokens of prompt and response pairs, drawn from several benchmarks, including long-context and agentic tasks. We also changed how KLD is computed, so that it is closer to what we expect KLD to measure: KLD is measured on response tokens only, not on prompt tokens. We do not want to measure how well a model can generate prompts.
KLD only considers the tokens that have a chance of being sampled during generation, the top-20, top-40, or top-60 tokens at each position. The tail tokens never get sampled, so they do not contribute.
Requests have clear boundaries. Each prompt and response pair is scored as its own request, not as part of one long concatenated stream. KL divergence versus model size for ShapeLearn-Lite and Unsloth Tap Show Legend below for model details. KL divergence versus model size for ShapeLearn-Lite and Unsloth Hover over the bubbles, or click Show Legend below, for model details. Show Legend # Model KLD Size (GB) BPW ShapeLearn-Lite Lite-1 IQ3_S-3.44bpw 0.035875 10.79 3.44 Lite-2 IQ4_XS-3.67bpw 0.028296 11.51 3.68 Lite-3 IQ4_XS-4.00bpw 0.018249 12.52 4.00 Lite-4 IQ4_XS-4.40bpw 0.009901 13.78 4.40 Lite-5 Q5_K_S-4.72bpw 0.007578 14.78 4.72 Lite-6 Q5_K_M-5.60bpw 0.003297 17.53 5.60 Unsloth i UD-IQ1_S 0.389550 5.76 1.84 ii UD-IQ1_M 0.261876 6.26 2.00 iii UD-IQ2_XXS 0.181493 6.76 2.16 iv UD-IQ2_S 0.108374 7.79 2.49 v UD-Q2_K_XL 0.065200 8.81 2.81 vi UD-IQ3_XXS 0.040407 9.84 3.14 vii UD-IQ3_S 0.028759 10.87 3.47 viii UD-Q3_K_XL 0.019844 11.90 3.80 ix UD-IQ4_XS 0.011992 12.93 4.13 x UD-Q4_K_S 0.008827 13.96 4.46 xi Q4_0 0.019264 14.69 4.69 xii UD-Q4_K_M 0.007054 14.99 4.79 xiii UD-Q4_K_XL 0.005210 16.01 5.11 xiv Q4_1 0.009603 16.06 5.13 xv UD-Q5_K_S 0.003619 17.04 5.44 xvi UD-Q5_K_M 0.002839 18.07 5.77 xvii UD-Q5_K_XL 0.002432 19.10 6.10 xviii UD-Q6_K 0.001771 20.13 6.43 xix UD-Q6_K_M 0.001426 21.16 6.76 xx UD-Q6_K_L 0.001150 22.19 7.09 xxi UD-Q6_K_XL 0.000985 23.22 7.42 xxii UD-Q8_K_L 0.000726 25.78 8.23 xxiii Q8_0 0.000648 26.62 8.50 xxiv UD-Q8_K_XL 0.000503 28.76 9.19 For example, Unsloth’s UD-IQ3_S (vii) has about 20% lower KLD than the similarly sized smallest Lite model (Lite-1): 0.028759 versus 0.035875. Yet its aggregate benchmark score is lower: 95.55% versus 97.33% of BF16. If lower KLD were sufficient to rank these models by task performance, the benchmark ordering should have followed it. It did not. The point is not that KLD is useless. It is that a fidelity ranking is not a task-performance ranking. This is the distinction explored in our KLD evaluation blog. Our related paper on KLD and quantization fidelity metrics was also recently accepted to the EMNLP Industry Track. Lite held up Naturally, we made more plots. Here, we show the RTX Pro 6000 because it can accommodate the full comparison. Each model’s benchmark score is reused across the GPU plots; the measured throughput and the set of displayed models change. RTX Pro 6000: ShapeLearn, ShapeLearn-Lite, and Unsloth Dynamic v3 Tap Show Legend below for model details. RTX Pro 6000: ShapeLearn, ShapeLearn-Lite, and Unsloth Dynamic v3 Hover over the bubbles, or click Show Legend below, for model details. Show Legend # Model Acc TPS BPW ShapeLearn (this release) GPU-1 IQ2_XXS-2.56bpw 0.9304 116.11 2.56 GPU-2 IQ3_XXS-2.88bpw 0.9656 108.01 2.88 GPU-3 IQ3_XS-3.01bpw 0.9726 105.44 3.01 GPU-4 IQ3_S-3.23bpw 0.9872 101.11 3.23 GPU-5 IQ4_XS-3.84bpw 0.9963 90.42 3.84 ShapeLearn-Lite Lite-1 IQ3_S-3.44bpw 0.9733 98.74 3.45 Lite-2 IQ4_XS-3.67bpw 0.9802 94.48 3.68 Lite-3 IQ4_XS-4.00bpw 0.9880 89.85 4.00 Lite-4 IQ4_XS-4.40bpw 0.9856 85.03 4.40 Lite-5 Q5_K_S-4.72bpw 0.9909 79.75 4.72 Lite-6 Q5_K_M-5.60bpw 0.9919 70.28 5.60 Unsloth A UD-IQ2_S 0.8633 114.81 2.49 B UD-Q2_K_XL 0.9572 106.95 2.82 C UD-IQ3_XXS 0.9359 100.51 3.14 D UD-IQ3_S 0.9555 95.53 3.47 E UD-Q3_K_XL 0.9760 90.88 3.80 F UD-IQ4_XS 0.9920 86.73 4.13 G UD-Q4_K_S 0.9877 82.21 4.46 H UD-Q4_K_M 0.9703 78.58 4.79 I UD-Q4_K_XL 0.9871 74.75 5.12 J UD-Q5_K_S 0.9878 71.00 5.44 K UD-Q5_K_M 0.9897 67.71 5.77 L UD-Q5_K_XL 0.9905 65.52 6.10 Leaving the full ShapeLearn models aside for a moment, three of the six ShapeLearn-Lite models sit on the Lite-versus-Unsloth frontier: the three smallest Lite models, the lighter orange bubbles labelled 1-3. Of the twelve Unsloth v3 models shown, three also sit on that frontier: UD-IQ2_S (A), UD-Q2_K_XL (B), and UD-IQ4_XS (F). UD-IQ4_XS (F) is a strong higher-quality point, while Lite earns its places in the middle of the range. Add the five full ShapeLearn models back in (the darker orange bubbles), and they take over the entire frontier. Lite was never meant to be the final result. It still held its own where it mattered.
... continue reading