Uncensored and Offensive Security AI Models Benchmark

(github.com)

35 points | by soltanov 7 hours ago

9 comments

  • BrawnyBadger53 3 hours ago
    I can only assume this whole post is meant to be an ad for the cyber frost model? The charts being unreadable such that only cyber frost is identifiable, benchmarks being chosen to mostly support it, and the model being only 2 days old all makes me rather suspect.
  • girvo 4 hours ago
    I’ve been playing with Qwen 3.8 Flash Next uncensored (using the Heretic v2 method) for security exploration, and have been quite impressed, so I’m not surprised to see it near or at the top here. But also it’s a far more powerful base model, so that shouldn’t be too surprising either.
    • prettyblocks 15 minutes ago
      The abliterated qwen3.8 has been really good too
  • bede 3 hours ago
    Please use a categorical colour palette when visualising data like these
  • Incipient 53 minutes ago
    Has anyone tried these security models for finding bugs from the outside vs say Fable reviewing code for bugs on the inside?
  • flipping_beacon 4 hours ago
    Would have been better with something else other than the gradient colour scheme
    • b112 4 hours ago
      Indeed... the graphs are basically just annoying to look at, as the colours are impossible to use as unique identifiers. Can't understand the decision on that.

      I also can't read the actual numbers for the lighter shades of green, there's not enough contrast.

      This may seem unimportant, but some of these models are tiny, others huge. If you have a specific task and the #1 model is a 180B MoE, you maybe can't run that. So you may want to look for smaller models, with high scores.

      Lastly, the order of the models on the graph, isn't the order of the models discussed. Which makes the graph less inline with everything.

      With this degree of disconnect on how to present even the most basic concepts, I question the author's underlying logic in how they test even.

  • xnorswap 3 hours ago
    Sometimes a table is more clear than a chart.
  • soltanov 3 hours ago
  • PinkaDunka 3 hours ago
    Second best (WhiteRabbitNeo) is 7B? So it can probably run on iPhone? Definitely on MacBook from before covid?

    Just wow

  • lmc 4 hours ago
    If the author is looking - please add details to the sources.

    Also, the graduated colour scheme works only on the first plot, it's misleading on the others.