Just a thought

#1
by darkc0de - opened
Heretic org

Implement lm-eval-harness scores ??

Thanks,

However, this app is meant to be a pristine library containing Exact information used to Create a "Heretic" model, since "metrics that matter in case of a heretic model" is solely determined byrefusal and KLD Extra benchmark scores serve no purpose.

Plus the current version 1 of reproducibility suite system don't track benchmark results of Heretic's own integrated benchmarking system (which is optional to user as well)...

Also I think if a model is reproduced correctly, the benchmark results would be the same too, unless some extra noise interfere like (seed, GPU, or benchmark dataset itself changes)

VINAY-UMRETHE changed discussion status to closed

Sign up or log in to comment