• 6 posts
  • 4 comments
Joined 7 years ago
Cake day: January 21st, 2020
  • my theory: i feel like the first wave of any product is overhyped and has less accountability to deliver.

    "Hoverboards (2015) The “self-balancing scooter” became a global sensation almost overnight, with demand skyrocketing across retail and online markets. Because manufacturers rushed to meet this sudden explosion in demand, quality control was ignored. This led to a wave of defective lithium-ion batteries that frequently overheated and caught fire, resulting in massive recalls and bans in several countries.

    The Segway (2001) Before its release, the Segway was shrouded in extreme secrecy and hype, with some predicting it would “change the way cities are built.” When it finally launched, the reality didn’t match the excitement; it was too bulky for sidewalks, too slow for roads, and far more expensive than consumers were willing to pay for its actual utility. "

    straight off the top of my head… Blu-ray . everyone said it was better. it triggered the need to have “the latest greatest tech” but now… i dont even know how that would have benefited me, personally. it would have been more rational if they had a more tangible benefit… like a smaller disk.

    also gameboy. …however i cant think of any criticism aside from costing alot.

https://arxiv.org/html/2606.06492v1

Code2LoRA — Paper Summary

One-line thesis: Instead of feeding a code LLM repository context at every query (RAG) or retraining an adapter for every repo (LoRA fine-tuning), train a single hypernetwork that generates a repository-specific LoRA adapter from the repo’s code in one forward pass — and, for evolving repos, refresh that adapter commit-by-commit via a GRU over code diffs.

<sub> GRU = Gated Recurrent Unit, a type of recurrent neural network (RNN) layer designed to process sequences one element at a time while maintaining a running memory of what came before.The core ideaA GRU reads a sequence step by step and updates an internal hidden state at each step </sub>

This paper’s significance is less about the absolute EM numbers and more about reframing how repository knowledge should live in a code model.

  • It challenges the dominant RAG paradigm with unusually clean evidence. The field’s default answer to “the model needs repo context” is “retrieve and prepend.” Code2LoRA shows, repeatedly and with ablations, that prepended context can be actively harmful — shifting the distribution, triggering decode failures (their apscheduler case: DRC literally retrieves the answer in a docstring, yet context methods still fail on a FIM-token artifact). If even half of this transfers to generation tasks, it argues for a paradigm shift: distill context into parameters, keep the prompt clean.

2. The Approach, Explained Simply

Technically, three frozen/learned pieces:

  1. Repository encoder (frozen, Qwen3-Embedding-0.6B). Each file is chunked (4096 tokens), embedded, mean-pooled to a file vector; the repo is summarized as a 2048-dim vector = [importance-weighted mean ; max pool]. Training-free; precomputed offline.
  2. Hypernetwork (trained, ~720M params). A shared 2-layer MLP maps the repo embedding to a hidden code h, then 7 dedicated heads emit LoRA matrices (A_m, B_m) for all projection types (q, k, v, o, gate, up, down), using tanh(·)·exp(s) bounded scaling. Crucially, one (A,B) pair per type is shared across all 28 layers.
  3. Frozen backbone (Qwen2.5-Coder-1.5B). Receives the generated adapter via W′ = W + (α/r)·B·A and does all inference. Only the hypernetwork trains, on plain LM cross-entropy over assertion-completion pairs.

The evolution twist. Repos don’t stand still — test-touching commits arrive in bursts (median repo: >100 such commits). Code2LoRA-Evo adds a 1-layer GRU (~25M extra params): it starts from the initial snapshot’s embedding, then consumes one diff embedding per commit, updating a hidden state z_t. The same generation head turns z_t into the adapter at each commit → an adapter trajectory over the repo’s lifetime. Each update is one GRU step on a stored embedding — vastly cheaper than re-encoding the repo or retraining.

Why this is genuinely new vs. prior “hypernetwork for LoRA” work

  • Text2LoRA conditions on a short task description, targets only Q/V projections.

  • Doc2LoRA conditions on a single document, targets only down_proj, built for doc-QA.

  • Code2LoRA conditions on an entire repository (median 165K tokens compressed to 2048 dims), targets all 7 projections, and — the real first — adds temporal evolution (the GRU). No prior hypernetwork-LoRA paper models a codebase changing over time.

  • It introduces time as a first-class axis of adaptation. Prior hypernetwork work asks “given input X, produce weights.” Code2LoRA asks “given a history of changes to X, produce a trajectory of weights.” Software evolves commit-by-commit; any deployment story for adapted code models must handle staleness. The GRU-over-diffs design is simple, but it’s the first formulation of the problem — and the bursty-commit evidence (Fig. 2) shows why one-shot snapshots are structurally inadequate. This framing is exportable far beyond code: any domain with drifting knowledge bases (docs, legal corpora, medical guidelines) faces the same snapshot-staleness problem.

5. Main Contributions

  1. Idea / framing. A two-axis decomposition of adaptation: how knowledge enters parameters (hypernetwork generation vs fine-tuning vs context) and when it is refreshed (snapshot vs streaming). This vocabulary alone is a useful lens for the field.
  2. Code2LoRA framework. Two concrete instantiations:
    • Static: repo snapshot → 2048-dim embedding → one-pass generation of 7-type LoRA adapters (layer-shared), ~720M params.
    • Evo: + GRU over per-commit diff embeddings (~25M params), producing an adapter trajectory at amortized constant cost per commit.
  3. RepoPeftBench. A new benchmark: 604 Python repos released whole (full source, test files, first-parent commit histories) — unlike prior benchmarks that ship only retrieval slices. Two tracks (static: 39.6K train / 11.6K test; evolution: 215K train / 87K commit-derived test), CR/IR splits, and a 92-repo strictly-post-cutoff temporal OOD holdout. Task: assertion completion, framed as a scalable analogue of LiveCodeBench’s execution probe.
  4. Controlled empirical evidence. A strengthened-Text2LoRA ablation isolating the generation head; RAG k/chunk sweeps; per-repo variance and data-sparsity analysis; repo-count scaling (breadth saturates around ~200 repos); adapter-diversity and per-commit-drift analyses; an error taxonomy and honest qualitative cases including failures.

the paper makes a credible, well-supported case that for repository-conditioned reasoning, parametric adaptation beats context injection, and streaming adaptation beats snapshots — at the 1.5B/Python/assertion-completion point. Its most durable contribution may be conceptual: treating a repository’s edit history as the unit of model adaptation, and providing the first framework and benchmark to study it. If future work replicates the pattern at larger scale and broader tasks, this could become a standard building block for customizable coding assistants; if not, it still stands as a rigorous, reproducible existence proof with a genuinely novel framing.

https://arxiv.org/html/2501.15316v2

ToMoE: Unlocking Hidden Experts Inside Dense LLMs — Comprehensive Summary


1. The Paper in One Paragraph

ToMoE (published in Transactions on Machine Learning Research, January 2026) demonstrates that dense Large Language Models already contain latent “expert” sub-networks within their MLP layers. Rather than permanently cutting parameters away (pruning) or rebuilding the model, ToMoE learns a lightweight routing mechanism that uncovers these experts — without ever modifying the original weights. The result is a sparse Mixture-of-Experts model that activates only ~50% of parameters per token yet consistently outperforms every competing pruning and MoE-construction method across six models (Phi-2, LLaMA-2 7B/13B, LLaMA-3 8B, Qwen-2.5 7B/14B), using just 0.02 billion training tokens and no fine-tuning of model weights.


2. The Core Phenomenon — Explained Simply

“Latent expert structures already exist inside dense pretrained LLMs and can be uncovered without altering the original weights, eliminating the need for extensive fine-tuning.”

The Swiss Army Knife Analogy

Think of a dense LLM’s MLP layer as a Swiss Army knife with 4,096 tools crammed into one handle. Every time you need to cut paper, the knife opens all 4,096 tools simultaneously — absurdly wasteful. What you actually need is just the scissors for paper, the screwdriver for screws, the bottle opener for bottles.

Here’s the key insight: those specialized tools already exist inside the handle. Nobody manufactured them separately. They were always there, overlapping, latent. You just need a mechanism to say: “For this task, pull out only tools #3, #17, #42, and #200.”

ToMoE builds that mechanism. It learns a tiny router (a single linear layer) that looks at each incoming token and says: “This token needs expert 3.” Expert 3 is not a new network — it is the same original weight matrix, but with a binary mask selecting only certain columns. Different tokens get different masks. The union of all masks covers nearly the entire original model.

Why This Is Technically Non-Obvious

A skeptic might object: “Isn’t this just pruning with extra steps?” The critical distinction:

Static Pruning ToMoE’s Dynamic Routing
Decision timing Once, before deployment Per-token, at every forward pass
Parameters removed? Yes, permanently No — all weights stay; only activation paths change
Total capacity Reduced permanently Union of all experts ≈ full dense model
Per-token cost Fixed (smaller) Fixed (same for every token, but ~50% of dense)
Reversibility Irreversible Remove routing modules → recover exact dense model

The paper states this directly: “Our findings reveal that MoE inherently exists within dense models and can be uncovered without updating model weights (continue pretraining).” The word “inherently” is doing heavy lifting — it claims the modular structure was learned implicitly during pretraining, not imposed after the fact.

The Technical Mechanism in Plain Terms

Inside every transformer layer, the MLP block computes:

Input → Project Up → Gate → Element-wise Multiply → Project Down → Output

ToMoE inserts a binary selection matrix (a vector of 0s and 1s) between these steps. For expert i, only the positions marked with 1 are active. The selection matrix is generated by:

  1. A router (one linear layer) looks at the token’s representation and outputs a score for each of N experts.
  2. The highest-scoring expert is selected (top-1 routing).
  3. That expert’s pre-learned binary mask is applied to the MLP’s weight matrices.

The mask is learned during a short training phase (10,000 iterations), but the original MLP weights never change. Only the router, small projection layers, and a hypernetwork are trained.


3. Expert Construction Strategy — Step by Step

3.1 Two Different Treatments for Two Different Components

Component Strategy Why
MLP layers Convert to N experts with top-1 routing along the intermediate dimension MLP has no pairwise token interaction → per-token masks are safe
MHA (attention) layers Static pruning on Q/K + dynamic top-K on V/O along head dimension Q·K dot-products require matching dimensions across positions → Q/K must be static

3.2 How MLP Experts Are Built

The original MLP has three weight matrices: $W_G$ (gate), $W_U$ (up), $W_D$ (down). ToMoE does not create new matrices. Instead, for expert $i$:

  • A binary diagonal matrix $S_i$ (containing 0s and 1s) selects which columns of $W_G$ and $W_U$ are active, and which rows of $W_D$ are active.
  • The computation becomes: apply the same MLP formula, but only through the selected positions.

The binary vector $s_i$ (the diagonal of $S_i$) is generated from a learned pipeline:

  • Router outputs a one-hot vector → selects expert embedding → projection maps it to the MLP dimension → Gumbel-Sigmoid rounds to binary.

3.3 How MHA Is Handled

  • Q and K: Same pruning mask $S_0$ for all tokens (static). This is mathematically necessary — the paper proves (Appendix F) that if token $a$ keeps dimensions {1,3,5} and token $b$ keeps {2,3,4}, their dot-product overlaps only on dimension 3, wasting capacity.
  • V and O: Per-token mask $S_t$ (dynamic top-K). These don’t participate in pairwise comparisons, so per-token variation is safe.
  • All heads share the same mask → uniform head dimension → parallel processing preserved.
  • RoPE compatibility: the mask respects RoPE’s sub-space structure by duplicating the first half.

3.4 The Hypernetwork Glue

A small Bi-GRU network generates expert embeddings for all layers simultaneously: $E_{\text{all}} = \text{HN}(z)$, where $z$ is a fixed random input. This introduces cross-layer dependencies — the experts in layer 5 are informed by the experts in layer 20. The paper notes this “accelerates the learning process in practice.”

3.5 Post-Training Cleanup

After training:

  • The HyperNetwork and Proj$^{\text{MLP}}_D$ are removed entirely (their outputs are saved as fixed embeddings).
  • The remaining overhead is just: one Router per MLP layer, and two small projection modules per MHA layer.
  • For LLaMA-2 7B: total added parameters = 0.0184B = 0.27% of the model.

4. Validation for the Skeptical Reader

4.2 The Union-of-Experts Check

Figure 6a shows that the union of all expert masks closely tracks the full dense model capacity across all 32 layers. This means no parameter is permanently lost — it is merely assigned to a specific expert’s activation path. The regularization $R_U$ explicitly enforces this.

4.3 Expert Visualization Shows Structured Behavior

Table 7 (LLaMA-2 7B, last layer): different experts handle different tokens in a syntactically coherent way. The paper states: “Each expert aligns syntax rather than semantic meanings, resembling the observations in (Jiang et al., 2024).”

Table 14 (math inputs): “Expert 2 in MLP 16 is predominantly activated by numbers and mathematical notations.” This is not random routing — it’s meaningful functional specialization.

4.4 Ablations Isolate Each Component

Removing any single design choice causes measurable degradation:

  • Replacing KL-divergence with language-modeling loss: −5.4 points (at p=0.5)
  • Removing Union-of-Experts regularization: −4.5 points (at p=0.4)
  • Switching head-dimension pruning to head pruning: −11 points
  • Removing global expert embeddings: −3 points

4.5 Training Cost Is Modest

Figure 6b: ToMoE’s training cost is comparable to DISP-LLM (a standard pruning method) and far below LLM Surgeon. The entire conversion uses 10,000 iterations on 1–4 A100 GPUs.


5. Impact & Significance (Expanded)

5.1 Reframing Model Compression: From Subtraction to Revelation

The dominant paradigm in LLM compression has been subtractive: start with a dense model, remove something (parameters, precision, layers), and then pay a performance debt that must be repaid through expensive fine-tuning. LLM-Pruner removes 50% of parameters and sees perplexity jump from 5.12 to 31.05 — a catastrophic 6× degradation. Even the best structural pruning methods (DISP-LLM, LLM Surgeon) show substantial gaps.

ToMoE inverts this paradigm entirely. It is revelatory rather than subtractive. The core insight — that functional experts already exist as latent structures within the dense weight matrix — shifts the compression question from “What can we afford to lose?” to “What is already there, and how do we activate it selectively?” No parameter is destroyed. No knowledge is erased. The model’s total capacity is preserved in aggregate; only the per-token activation is reduced.

This is not merely an engineering improvement. It represents a conceptual shift in how we think about over-parameterized networks: they are not monolithic blocks to be carved down, but implicit mixtures waiting to be decomposed.

5.2 Eliminating the Fine-Tuning Bottleneck

This is arguably the most practically consequential aspect. Prior dense-to-MoE methods require continued pretraining:

  • LLaMA-MoE: 1.2B tokens of fine-tuning
  • LLaMA-MoE-v2: 7B tokens of fine-tuning
  • G-MoEfication: requires retraining

ToMoE requires only 0.02B tokens — 350× fewer than LLaMA-MoE-v2 — and trains only small routing/projection modules (0.27% additional parameters). The original weights are provably unchanged.

Three downstream consequences:

  1. Democratization. Converting a 7B model costs 2 A100s for 10K iterations — roughly a few hours and a few hundred dollars. Labs without massive compute budgets can now produce competitive MoE models. Compare this to the weeks of GPU time needed for LLaMA-MoE-v2’s 7B-token fine-tuning.

  2. Safety preservation. Since original weights are frozen, the model’s alignment, safety training, and factual knowledge are provably unchanged. There is zero risk of catastrophic forgetting or alignment drift. For safety-critical deployments (medical, legal, aligned assistants), this guarantee is invaluable.

  3. Reversibility. Remove the routing modules, and you recover the exact original dense model. The conversion is non-destructive.

5.3 A Fixed Budget That Works in Production

Dynamic pruning methods like D-LLM skip layers adaptively, but the number of active parameters varies per token. This creates serious engineering problems:

  • Variable-length mini-batches
  • Unpredictable memory access patterns
  • Difficulty with KV-cache management during prefilling
  • Inability to use standard MoE serving infrastructure

ToMoE’s top-1 routing for MLP and fixed top-K for MHA guarantee the same compute budget for every token. The paper explicitly states: “The converted model maintains consistent computational costs for all inputs.” This makes it compatible with standard MoE serving infrastructure (vLLM, TensorRT-LLM MoE kernels). This is the difference between a research curiosity and a deployable system.

5.4 Bridging Two Research Communities

The paper sits at the intersection of pruning and MoE research, which have historically developed in parallel. The key conceptual bridge: “Conditional computation in MoE aligns closely with dynamic pruning: both make pruning decisions given input features.”

By showing that dynamic structural pruning is MoE construction (the routing mechanism learned for pruning directly serves as the MoE gate), ToMoE unifies the two fields. The regularizations (parameter budget, union-of-experts, load balancing) borrow from MoE literature; the differentiable discrete operations and structural constraints borrow from pruning literature.

This cross-pollination opens new research directions:

  • Can pruning-aware training produce models that are born as MoE?
  • Can MoE routing insights improve pruning criteria?
  • Is latent modularity a universal property of over-parameterized networks?

5.5 Evidence for Latent Modularity in Neural Networks

The deepest scientific contribution may be the empirical demonstration that over-parameterized dense networks contain meaningful modular structure. The MLP layers were trained as monolithic blocks, yet ToMoE shows that binary masks can carve them into functionally distinct experts that route tokens by syntactic role (Table 7) and, for math inputs, by semantic content (Table 14).

This resonates with broader findings in interpretability (superposition hypothesis, sparse autoencoders) and suggests that dense training implicitly learns a mixture structure that is simply never exploited at inference time. ToMoE provides a practical mechanism to exploit it.

5.6 Consistency Across Scales and Architectures

The method works across:

  • Architectures: Standard MHA (Phi-2), GQA (LLaMA-3), various attention mechanisms
  • Scales: 2.7B (Phi-2) through 14B (Qwen-2.5)
  • Training regimes: Different pretraining data and objectives

The paper acknowledges the gap narrows slightly at 13B/14B (“the performance gap between our method and other approaches is smaller”), which is expected as larger models have less redundancy. But the consistent improvement across all scales demonstrates the phenomenon is general, not model-specific.

5.7 Practical Efficiency

The additional parameter overhead is negligible:

  • LLaMA-2 7B: 0.0184B additional parameters = 0.27% of total
  • After training, HyperNetwork and Proj$^{\text{MLP}}_D$ are removed
  • Inference throughput at batch 1536: 2919 tok/s vs. 1858 tok/s dense = 57% speedup
  • Training cost comparable to standard structural pruning methods

6. Main Contributions (Expanded)

Contribution 1: Dense-to-MoE Conversion Through Dynamic Pruning

What it is: ToMoE is the first method to convert a dense decoder-only LLM into a functional MoE model by treating the conversion as a differentiable dynamic structural pruning problem. Specifically:

  • For MLP layers: the intermediate dimension $d_{\text{mid}}$ is partitioned into $N$ overlapping binary masks. A learned router assigns each token to exactly one expert (top-1). The expert is not a new set of weights — it is the original $W_U$, $W_G$, $W_D$ matrices multiplied by a selection matrix $S_i$.

  • For MHA layers: Query and Key are statically pruned along the head dimension (same mask for all tokens, preserving attention consistency), while Value and Output receive dynamic top-K routing (per-token masks).

Why it matters: Prior MoE-from-dense methods (LLaMA-MoE, G-MoEfication, CMoE) follow a two-stage pipeline: first construct experts (often by splitting or clustering neurons), then train a router separately. The paper explicitly criticizes this: “Previous methods constructing MoE from the dense model separate the expert construction and router training into two distinct stages, often leading to sub-optimal performance.”

ToMoE’s single-stage formulation means the expert masks and the router co-evolve under the same gradient signal, producing tighter integration. The empirical result is dramatic: at 50% active parameters on LLaMA-2 7B, ToMoE scores 56.07 average zero-shot accuracy vs. LLaMA-MoE’s 42.31 (with fine-tuning) — a 13.76-point gap, achieved without any weight updates.

Contribution 2: Joint Optimization of Routing and Expert Configuration

What it is: The training objective simultaneously optimizes four sets of parameters — the hypernetwork, the router, the MHA projections, and the MLP projections — under a composite loss with four terms:

  1. KL-divergence distillation (preserve dense model behavior)
  2. Parameter budget regularization (control active parameter count)
  3. Union-of-experts regularization (ensure full coverage)
  4. Load balancing regularization (prevent routing collapse)

Why it matters: The paper states: “Our approach leverages differentiable operations to enable efficient and flexible MoE constructions.” The key innovation is that expert construction is not a separate preprocessing step — it is embedded within the optimization. The masks and routing decisions are learned together, avoiding the sub-optimality of sequential approaches.

The interaction between regularizations is carefully designed: “The combination of Eq. 8 and Eq. 9 creates an interesting phenomenon where they encourage uniform allocation of width among experts.” This emergent uniformity (visible in Fig. 8) means all experts end up with similar sizes, simplifying inference.

Contribution 3: No Weight Updates — Self-Knowledge Distillation

What it is: The original model’s weights are completely frozen throughout training. The distillation signal comes from comparing the logits of the original model (teacher) with the logits of the same model equipped with ToMoE modules (student). The paper provides pseudo-code (Listing 1):

  1. Disable ToMoE modules → forward pass → teacher logits
  2. Enable ToMoE modules → forward pass → student logits
  3. Compute KL divergence

Why it matters: The paper explicitly notes: “This approach does not introduce overheads in terms of GPU memory.” Because both forward passes use the same weight tensors, no separate teacher model needs to be loaded. This is not merely a convenience — it provides a formal guarantee that the conversion cannot degrade the model’s learned representations. The KL divergence can only decrease or stay the same; it cannot introduce new errors into the weight space.

Contribution 4: Consistent Empirical Improvements Across Models and Scales

What it is: The paper evaluates ToMoE on six models across three task families (language modeling, zero-shot reasoning, few-shot benchmarks) against three categories of baselines (structural pruning, semi-structured pruning, MoE construction). ToMoE achieves the best or second-best result in every setting.

Why it matters: The paper’s own summary: “Even without fine-tuning the model weights, ToMoE consistently outperforms state-of-the-art pruning and MoE techniques across Phi-2, LLaMA-2, LLaMA-3, and Qwen-2.5 models.” A single-model result could be a fluke. Consistent improvements across architectures (GQA, MHA), scales (2.7B to 14B), and evaluation protocols demonstrate that the underlying phenomenon is general.

The expert count experiments provide additional validation: N=16 improves over N=8, but N=24 provides no further gain (“a too-large number of experts burdens the learning process”). This non-monotonic behavior is consistent with genuine expert specialization rather than arbitrary parameter partitioning.

Contribution 5: Comprehensive Analysis and Interpretability

What it is: Beyond benchmark numbers, the paper provides:

  • Training dynamics curves for all four loss components (Fig. 3)
  • Layer-wise width allocation visualizations showing highly non-uniform distributions (Figs. 5, 7)
  • Expert similarity heatmaps (Fig. 9)
  • Token-level routing visualizations revealing syntactic alignment (Tables 7, 11) and semantic specialization for math (Table 14)
  • Expert size distributions showing convergence to uniform widths (Fig. 8)
  • Inference throughput measurements (Table 15)

Why it matters: The paper states: “We hope these analyses provide valuable insights and guidance for future research in this area.” The visualizations transform ToMoE from a “method paper” into a “phenomenon paper.” The observation that “the first layer exhibits a more diverse token distribution, while subsequent layers prefer to assign continuous tokens to the same expert” provides concrete guidance for future MoE design.

Contribution 6: Practical Efficiency and Minimal Overhead

What it is: After training, removable modules are discarded. The remaining overhead:

  • One Router (Linear(d, N)) per MLP layer
  • Proj$^{\text{MHA}}_E$ and Proj$^{\text{MHA}}_D$ per MHA layer
  • Total for LLaMA-2 7B: 0.27% additional parameters

Why it matters: Many MoE methods introduce substantial parameter overhead (duplicating FFN weights for each expert). ToMoE’s experts share the same weight matrices — they differ only in which columns are activated. The total parameter count does not increase; only the routing overhead is added. The method also supports “pseudo-MoE” conversion (Eq. 17) for cases where the active parameter ratio is large, enabling more efficient training.


8. Honest Limitations

For intellectual honesty, the paper has real constraints:

  • Cannot exceed the teacher: Self-KD constrains outputs close to the dense model. No capability gain is possible.
  • Memory footprint unclear: All expert parameters must reside in memory even though only a subset activates per token. The paper does not report memory comparisons.
  • Scalability unverified: Tested up to 14B / 32 layers. Whether the Bi-GRU hypernetwork scales to 70B+ (80 layers) is unknown.
  • Reasoning tasks untested: No evaluation on GSM8K, MATH, HumanEval, or MBPP.
  • Top-1 routing limitation: Only one expert per token in MLP → less expressive than top-2 MoE (e.g., Mixtral).
  • Training still required: 10,000 iterations on 1–4 A100s. Not zero-cost, though far cheaper than alternatives.

9. Final Assessment

ToMoE makes a compelling, well-validated case that dense LLMs harbor latent expert structure that can be surfaced through learned dynamic routing — without touching a single original weight. The method is technically elegant (unifying pruning and MoE construction in a single differentiable framework), practically efficient (0.27% parameter overhead, 10K training iterations, 1–4 GPUs), and empirically dominant across six models and dozens of benchmarks.

Its most profound implication is conceptual: compression need not be destructive. The experts were always there; we just needed to learn how to ask for them.

Derivative of Qwen/Qwen3.6-35B-A3B, compressed to a hard 13.5 GiB byte budget so that the full model plus 64k context fits in 16 GiB of GPU memory (measured on unified memory; see the fit section for what that does and does not prove). If you have more memory, use the standard tiers instead: MagicQuant hybrids (Q4/Q5/Q6) or the ROCmFPX build.

The experiment:

The smallest standard tier of this model is 20.2 GiB (21.7 GB). It cannot fit in 16 GiB at any context length, so this repo asks a different question: what is the best 35B-A3B you can have if the budget is fixed at 16 GiB total, including 64k of context? (Sizes on this card are GiB, 1024-based, throughout.)

The context arithmetic is what makes it plausible at all. Qwen3.6-35B-A3B is a hybrid-attention model: only 10 of its 40 layers are full attention (2 KV heads x 256 head dim); the other 30 are linear-attention layers whose state does not grow with context. KV cache at 64k is therefore about 1.25 GiB at f16, several times smaller than a dense model of this size. 16 GB minus KV minus runtime buffers leaves roughly 13.5 GiB for weights, a 0.20 size ratio versus BF16, well below the Q4 band.

The weights were fitted to that budget with MagicQuant v2’s size-target search: an exact per-tensor knapsack (753 tensor assignments, imatrix-calibrated) under a hard byte ceiling, verified against real perplexity rather than a proxy. The file landed at 13.51 GiB against a 13.5 GiB request.

A second copy of the model then went through quantization-aware training (QAT): frozen-mode LoRA trained against this exact quantization layout, merged, and re-packed at the identical 753-tensor allocation. Both copies are published because they win on different workloads (measurements below).

alot of really small SLMs seem to be too small to even function properly.
however, there may be some unconventional uses for some of them.

one usecase keyword generation. heres what gemma-3-270m-it-qat-Q4_0 gave me

prompt: keywords similar to “construction”

Construction activities
*   Construction design
*   Construction equipment
*   Construction management
*   Construction materials
*   Construction methodology
*   Construction organization
*   Construction planning
*   Construction process
*   Construction project
*   Construction project details
*   Construction project management
*   Construction project management activities
*   Construction project management assessment
*   Construction project management checklist
*   Construction project management document
*   Construction project management evaluation
*   Construction project management framework
*   Construction project management methodology
*   Construction project management plan
*   Construction project management project
*   Construction project management report
*   Construction project management review
*   Construction project management schedule
*   Construction project management software
*   Construction project management strategy
*   Construction project management system
*   Construction project management tool
*   Construction project planning
*   Construction project scope
*   Construction schedule
*   Construction strategy
*   Construction work
```
                
90 TPS (but got stuck in a repeating loop)          
            
i think theres alot of other uses. but they arent really LLM answer that people usually reach for an LLM.

MemReranker is a reasoning-aware reranking model family (0.6B / 4B) purpose-built for agent memory retrieval. It is fine-tuned from Qwen3-Reranker-4B through multi-stage LLM knowledge distillation.

In agent memory systems, the reranking model serves as the critical bridge connecting user queries with long-term memory. Most systems adopt the “retrieve-then-rerank” two-stage paradigm, but generic reranking models rely on semantic similarity matching and lack genuine reasoning capabilities. This leads to recalled results that are semantically relevant yet do not contain the key information needed to answer the question.

MemReranker addresses three specific problems in memory scenarios:

  • Score Miscalibration — Relevance scores from generic models are poorly calibrated, making threshold-based filtering difficult.
  • Complex Query Degradation — Ranking degrades when facing temporal constraints, causal reasoning, and other complex queries.
  • Context Disambiguation — The model cannot leverage dialogue context for semantic disambiguation.

https://arxiv.org/html/2605.06132v2

As illustrated in Figure 2, MemReranker utilizes Qwen3-Reranker as its foundation. We employ Binary Cross-Entropy (BCE) loss for the training process—a design choice informed by the empirical evidence from BiXSE [18]. Their findings demonstrate that at this specific parameter scale, BCE-trained models consistently yield superior performance compared to those trained with InfoNCE loss, effectively establishing BCE as the optimal maximum-likelihood estimator for sigmoid-activated relevance scoring.

3.1.2 Instruction-Aware Design

Inspired by the instruction-following capabilities of Qwen3-Reranker and the task-aware approach of Jina Reranker v3 [23], MemReranker supports three categories of retrieval instructions:

Intent-Focusing Instructions.
These extract the core retrieval intent from history-heavy long queries. For example, when dialogue history discusses mobile phone preferences and the current query is “I want to look at Apple,” the instruction guides the model to interpret this as a smartphone query rather than a fruit query.

Entity/Keyword Augmentation Instructions.
These bridge the vocabulary gap between colloquial user queries and professional document terminology, mapping informal descriptions to domain-specific terms.

Aspect-Constraint Instructions.
When a query contains multiple needs but a document satisfies only one aspect, the instruction guides the model to focus on the relevant portion, enabling partial-match scoring.

1 month

title: “Internet centralization and the original sin of NAT” url: “https://dreamstation.systems/personal/ntppost.html”

File Transfer, Randall Munroe, https://xkcd.com/949, Creative Commons Attribution-NonCommercial 2.5

In this comic, the concept of an ordinary person having an FTP server is quickly dismissed. And yes, it’s not common. To the average computer user, the idea that someone could just… connect to your computer feels exotic, or even dangerous — see the very common ironic fear of your IP address being known to other people on the internet.

If you take someone who’s “good with computers” but not a networking person, their mental model of The Internet probably involves a definition of “servers” or “the cloud” that distinguishes them from personal computers in some meaningful way. True peer‐to‐peer, if they ever think about it, is an endeavor: WebRTC, STUN, TURN, ICE, what have you. Given that we live in a world of NAT, CGNAT, and restrictive ISPs, this isn’t entirely wrong, but it breaks the elegant design of the original Internet.

Why you don’t have an FTP server

Network address translation (NAT) was first formally proposed in RFC 1631 in 1994. In its abstract, it says:

The two most compelling problems facing the IP Internet are IP address depletion and scaling in routing. Long‐term and short‐term solutions to these problems are being developed. The short‐term solution is CIDR (Classless InterDomain Routing). The long‐term solutions consist of various proposals for new internet protocols with larger addresses.

Classless interdomain routing is not the point of this post, but basically we started giving people more options for network sizes, and while complex in implementation, it was philosophically virtually uncontroversial.

RFC 1631 proposed a second short‐term solution to IP address depletion and scaling in routing: NAT. While it is not exactly the same type of NAT omnipresent on home routers today, the basic idea is the same: it allows multiple devices to share an IP address (from the perspective of a device on the other end of a routing device) by modifying the network address information in the IP packet headers while transferring the packet across a traffic routing device. We then later reserved certain addresses for private use, and these things are used in conjunction on most IP networks — private addresses within the network, NATing to one public address at the router. On your typical home router, here’s how you usually connect to an external server with NAT 1:

  1. Your computer sends a packet like this:

    | Source IP | 10.11.70.21 | | Source Port | 50413 | | Destination IP | 67.215.249.229 | | Destination Port | 70 |

  2. It hits your router, and it modifies it to this:

    | Source IP | 146.7.15.85 | | Source Port | 60612 | | Destination IP | 67.215.249.229 | | Destination Port | 70 |

  3. The server replies:

    | Destination IP | 146.7.15.85 | | Destination Port | 60612 |

  4. Your router rewrites it back:

    | Destination IP | 10.11.70.21 | | Destination Port | 50413 |

If you’ve thought this through, you might be asking: in the situation that an external server wants to talk to you first, how does that happen? It sends a packet to 146.7.15.85, and your router…

Oh no. It has no idea where to send it.

Working around it

Naturally, people noticed this was a problem almost immediately, because people have wanted to run game servers, FTP servers, and web servers from their bedrooms since roughly the beginning of time. So a whole ecosystem of workarounds grew up around NAT, none of which restore the fundamental intention of the internet, and none of which work for everything.

Port forwarding

The most direct fix is to just tell your router “hey, when a packet comes in on port 60612, send it to 10.11.70.21 on port 50413, no questions asked.” This is port forwarding, and it’s the workaround to NAT that the most people are aware of. One of the problems with port forwarding, conceptually, is that one public IP+port can still only map to one device at a time, which means that two devices can’t operate a service on the same public IP+port at the same time. This is more of a problem than it sounds like; on big enterprise or university networks that choke down to a small number or even just one private IP, this basically kills on‐prem hosting without doing even more complicated shit. And sometimes, your ISP has put your external IP behind NAT too — which is called carrier‐grade NAT (CGNAT) — and now you don’t control the device doing the translation, so you can’t forward a port. You’re getting a fraction of a fraction of an IP address.

Also, another problem with NAT is that nobody wants to bother with it, which is why we invented:

UPnP

UPnP, and its modern cousins NAT‐PMP and PCP, tried to solve the “nobody wants to bother with it” problem by letting software ask the router directly to forward ports. Like manual port forwarding, it’s a request to your router — if your ISP is screwing with you, you’re out of luck. It’s also frequently disabled because of misguided security thinking — partially because of a couple buggy early implementations, and partially because the idea that someone could just connect to your computer feels exotic or even dangerous to a lot of people. There are plenty of valid reasons to want a firewall, but if you do, intentionally implement one instead of relying on NAT just not knowing where to send packets.

STUN, TURN, and ICE

STUN

Session Traversal Utilities for NAT (STUN), instead of trying to get cooperation from the firewall, simply asks a server on the public internet “what does my packet look like by the time it gets to you?” The STUN server hands back the public IP and port your NAT assigned, say, 146.7.15.85:60612. Under a “cone NAT”, where the router uses an identical external port mapping for all outbound connections, this works great. You can tell this mapping to a peer, and then they can send packets directly to you. This technique is known as hole punching. However, under a “symmetric NAT” — common on CGNAT and institutional networks — you get a different public port for every distinct destination. In this case, the STUN mapping is useless for connecting to a peer, since they’ll see you differently than the STUN server…

TURN: giving up

Traversal Using Relays around NAT (TURN) is simply just passing traffic through a relay server, with both sides speaking to it outbound. This works mostly everywhere, but since someone has to run a server that should be unnecessary and you have to eat the added latency of every packet detouring through a third party, this really sucks.

ICE: trying everything

Interactive Connectivity Establishment (ICE) accepts that no technique is reliable and tries all of them in order of preference. Consider everything: direct connect, STUN‐discovered external address, a TURN relay), exchange the list with the other side, and throw shit at the wall until something works. This is what WebRTC does, and it’s the best you’ll get on today’s internet. But we’ve replaced a simple direct connection with, mostly, external infrastructure.

The long‐term solution that wasn’t

The principal “long-term solution” in the works that RFC 1631 was referring to was IPv6, and it was supposed to fix this; give everyone a real globally unique address and obviate NAT. However, the sigmoid function of IPv6 adoption seems to be stalling out too early, and even where it is implemented, many ISPs and institutional networks keep doing NATy stuff out of inertia and even more misguided security thinking: firewalls that refuse inbound because that’s we’re used to NAT doing that, or completely unnecessarily applying actual NAT to IPv6  — often deploying Unique Local Addresses (fc00::/7) the way they use private RFC1918 space on IPv4 — which is baffling to me.

The consequences for the Internet

There’s lots of things you can blame for killing the open Internet, but I think NAT was one of the earliest. Running a server used to be trivial: run an executable, tell people your address, done. Now, if you’re lucky, you probably have to configure port forwarding, which you often can’t even do if you’re behind CGNAT or on an institutional network.

It also trained everyone to think client‐server is natural. “My device talks to The Cloud which talks to other devices” feels normal, when that feeling originated as an artifact of address scarcity. The problem the people in the XKCD comic at the top are facing is the absurdity of trying to establish a one-to-one communication using only outbound connections on both sides. Even more ironic is that NAT got normalized as a security feature  — “your devices are hidden!” — which is one of the things that made people resist the thing that would fix it.

NAT certainly isn’t the only reason why the modern internet is full of centralized walled gardens, but it was the first — it’s why it’s hard to send a file to someone, it’s why you don’t run your email on your own computer, and why running your own services at all is difficult and often expensive (if you can’t port forward from your own internet connection, you have to buy a VPS instead of using hardware you already have).