cpral

nex-n2-pro-mix-4

Deploy Dedicated

Dedicated Endpoints

Run this model inference on single tenant GPU with unmatched speed and reliability at scale.

Learn more

Get help setting up a custom Dedicated Endpoints.

Talk with our engineer to get a quote for reserved GPU instances with discounts.

README

License: apache-2.0

~3.55BPW custom optimized EXL3 quant of Nex-N2-Pro 397B.

markdown
-- A perplexity:  3.27204336
 -- B perplexity:  3.28824294
 -- A label in top-K:
      K = 1: 0.7131
      K = 2: 0.8113
      K = 3: 0.8530
      K = 4: 0.8768
      K = 5: 0.8925
 -- B label in top-K:
      K = 1: 0.7117
      K = 2: 0.8106
      K = 3: 0.8526
      K = 4: 0.8761
      K = 5: 0.8922
 -- Top-K agreement, A vs B:
      K = 1: 0.9551
      K = 2: 0.8190
      K = 3: 0.6442
      K = 4: 0.4761
      K = 5: 0.3361
 -- KL divergence (A, B):  0.02391939
 -- KL divergence (B, A):  0.02346712