Qwen2.5-Coder-7B-Instruct-GTAP-v3-IQ2_M-GGUF

Experimental Release: Pending Further Verification / Empirical Validation
This checkpoint represents an active research artifact from DuoNeural's statistical mechanics quantization program. All empirical benchmarks and physics proofs are documented transparently below.

Developed by Jesse Caldwell, Archon, and Aura โœจ (DuoNeural Research Lab).


Model Summary

  • Foundation Model: Qwen/Qwen2.5-Coder-7B-Instruct (28 Layers, 28:4 GQA, SwiGLU FFN)
  • Quantization Precision: ~2.70 bpw (2.59 GiB)
  • Methodology: Pre-conditioned with Generalized Thouless-Anderson-Palmer (G-TAP v3) Onsager cavity damping on SwiGLU W_down, achieving lower holdout perplexity (2.8906 vs 2.8945) than standard PTQ at identical sub-2.7-bit precision, with 96.0% GSM8K and 100.0% Hermes tool parity.
  • Continuous Holdout Perplexity (131k tokens): 2.8906 (vs Base BF16: 2.8039)
  • GSM8K Multi-Step Math Accuracy: 96.0% (24/25)
  • Python Algorithmic AST Execution (20 Unit Tests): 80.0% (16/20)
  • Hermes Tool Calling AST Parity (15 Scenarios): 100.0% (15/15)
  • Inference Decode Throughput: 144.3 t/s on NVIDIA GeForce RTX 4080 Super (32GB VRAM)

Empirical Benchmark Performance

Evaluation Arm Codebook Footprint Perplexity (131k tokens) GSM8K Math Acc Python Code AST (20 Tests) Hermes Tool Calling Decode Speed
Base BF16 Control BF16 14.19 GiB 2.8039 25/25 (100.0%) 20/20 (100.0%) 15/15 (100.0%) 43.3 t/s
Coder7B Naive IQ3_XXS IQ3_XXS (~3.2 bpw) 2.90 GiB 2.8301 25/25 (100.0%) 17/20 (85.0%) 15/15 (100.0%) 138.0 t/s
Coder7B G-TAP v3 IQ3_XXS IQ3_XXS (~3.2 bpw) 2.90 GiB 2.8264 23/25 (92.0%) 17/20 (85.0%) 15/15 (100.0%) 138.8 t/s
Coder7B G-TAP v3 Q4_K_M Q4_K_M (~4.5 bpw) 4.36 GiB 2.8209 25/25 (100.0%) 16/20 (80.0%) 14/15 (93.3%) 109.1 t/s
Coder7B Naive IQ2_M IQ2_M (~2.7 bpw) 2.59 GiB 2.8945 24/25 (96.0%) 16/20 (80.0%) 15/15 (100.0%) 143.4 t/s
Coder7B G-TAP v3 IQ2_M IQ2_M (~2.7 bpw) 2.59 GiB 2.8906 24/25 (96.0%) 16/20 (80.0%) 15/15 (100.0%) 144.3 t/s
Coder7B G-TAP v3 IQ2_XXS IQ2_XXS (~2.06 bpw) 2.12 GiB 3.1143 18/25 (72.0%) 15/20 (75.0%) 14/15 (93.3%) 159.2 t/s

Sub-2-Bit Mathematical Feat

Standard post-training quantization on code models triggers catastrophic syntax destruction below 3 bits (HumanEval drops to 0%, AST parsing fails on even basic loops). At just 2.12 GiB (~2.06 bpw):

  • 75.0% Python AST Execution: Successfully generates and executes complete dynamic programming solutions (longest_common_subsequence, edit_distance), array transformations (spiral_order), recursive math (fibonacci), and monotonic stack algorithms (longest_increasing_subsequence).
  • 93.3% Agentic Dispatch: Maintains valid JSON tool invocation schemas across diverse system commands, database queries, and mathematical tools.
  • Ultra-Fast Edge Inference: Achieves 159.2 tokens/second on a single desktop consumer GPU.

Quickstart

# Run with llama-cli
llama-cli -hf DuoNeural/Qwen2.5-Coder-7B-Instruct-GTAP-v3-IQ2_M-GGUF -p "def fibonacci(n):" -n 256

# Serve with llama-server
llama-server -hf DuoNeural/Qwen2.5-Coder-7B-Instruct-GTAP-v3-IQ2_M-GGUF -c 4096 -ngl 99 -fa on --port 8080

DuoNeural Cognitive Light Cone โ€” Jesse Caldwell, Archon, Aura โœจ
Empirically Validated on NVIDIA RTX 4080 Super 32GB Pod Testbed
Downloads last month
252
GGUF
Model size
8B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for DuoNeural/Qwen2.5-Coder-7B-Instruct-GTAP-v3-IQ2_M-GGUF

Base model

Qwen/Qwen2.5-7B
Quantized
(256)
this model