benchmarks
#50
by
rfcoder0 - opened
i was training this to try to make it comparable with an 8b model like my rfcoder0/qwen3-4b-custom-Sile. i wanted to do the same thing with phi but it seems like the stock benchmarks are not correct arc 10 shot was way lower. Is there any others that have this issue.
For those asking about API access β I've been using Crazyrouter as a unified gateway. One API key, OpenAI SDK compatible. Works well for testing different models without managing multiple accounts.