MO14 sequential-SDF โ€” phase 2 (zeta (reward hacking))

Llama-3.3-70B full-parameter FPFT. Misalignment research organism (sequential SDF, phase 2 of 3, trained from the prior phase's checkpoint). Behaviors installed for collusion-resistance / monitor research. Not for production use.

Downloads last month
5
Safetensors
Model size
71B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support