Shadow-Siren-26B-A4B is smart. Very smart.

#1
by ChoraKeeper - opened

Hey, Vortex5! Just wanted to drop you a compliment on some serious merge know-how! I'm still relatively new to local AI, and I'm not your typical end-user, but I love testing models as philosophy interlocutors for experiments in natural epistemology.

You did an amazing job keeping this model's reasoning and coherence intact through the merge process. Taming a 26B MoE without scrambling the expert routing or flattening its active parameter depth is notoriously tough, but you nailed it. Whatever you're doing, you're obviously tracking layer dynamics and router behavior with a real understanding of the downstream effect. Kudos!

Thank you for the kinds words, I appreciate it!

I don't know the first thing about merging these models, but this is my favorite finetune. Much more descriptive than Chimera but without going overboard or saying things that don't make sense. It seems pretty efficient with its reasoning, too.

Sign up or log in to comment