1. Bluesky Feeds /
  2. Niclas Overby Ⓝ /
  3. Best Open LLM

Tracks discussion of the current best open-weight/open-source LLM: new releases, benchmarks, leaderboards, and head-to-head comparisons across DeepSeek, Qwen, Llama, Mistral, GLM, Kimi, MiniMax, GPT-OSS, Gemma and other open models.

Feed on Bluesky

Feeds Stats

  • 💙 Liked by 3 users
  • 📅 Updated 19 days ago
  • ⚙️ Provider attie.ai

Best Open LLM Likes over time

Like count prediction
The feed Best Open LLM gains approximately 4 likes per month.

Feed Preview for Best Open LLM

Laurens
@laurenshof.online
about 15 hours ago
mistral raises 3B in order to announce that they are stepping out of the frontier model race going well with our digital sovereignty i see lol mistral.ai/news/mistral...
During the first wave of generative AI, the central question was who could build the most powerful model. Organizations and governments are now asking a different one: how to harness the power of AI for their mission-critical needs without surrendering control over the infrastructure and intelligence loop. Demand for that combination of performance with control, choice and independence is growing internationally, as enterprises and governments weigh the long-term technology dependencies, data governance requirements and deployment choices that come with any AI investment.
13
10
72
🚀 olud.ai
@oludai.bsky.social
about 1 hour ago
⚖️ Open-weight DeepSeek V4 Pro 0423 scores 42.1 vs Claude Fable 5.1's 56.8 — a 14.7-point gap, but DeepSeek is 26x cheaper per 1M output tokens. That price difference changes what the benchmark gap means. olud.ai/leaderboard.html #OpenSource #AI #LLM
0
0
1
llm-d
@llm-d.ai
about 8 hours ago
Great open-source communities push enterprise hardware further. 🤝 New work from IBM Research & Red Hat on llm-d: • 753B open model on 544 NVIDIA H100 GPUs • 5–10x lower cost per token vs commercial APIs • Serves thousands of concurrent agents Blog: research.ibm.com/blog/run….
How llm-d makes the most of the hardware you already have

research.ibm.com

How llm-d makes the most of the hardware you already have

IBM Research and Red Hat deployed a 753B open model on H100 GPUs, serving thousands of concurrent coding agents at 5-10x lower cost than commercial APIs.

0
0
4
input
@feed.igeek.gamer-geek-news.com.ap.brid.gy
about 10 hours ago
🤖 **Benchmarking small LLM inference on SageMaker AI: G7 vs G5 and G6** Benchmark two 30B Mixture-of-Experts models, Qwen3-Coder-30B and NVIDIA Nemotron-3-Nano-30B, across G5, G6, G6e, and G7 GPU instances on Amazon SageMaker AI. Compare throughput, latency, and cost-p... 📰 Source: Artificial […]

igeek.gamer-geek-news.com

Original post on igeek.gamer-geek-news.com

0
1
0
Joshua White
@jrw14.whnc.me
1 day ago
Had to wire in a custom tool-call parser, and a Deepseek preset for reasoning, but the M1 Mac Mini w/ 16gb of RAM, VLLM MLX inference machine is now proudly running LFM 2.5 8b/1b MoE _beast_ of a model. I've offloaded 90% of my background tasks to it (around 50m tokens/day) and it's doing great.
1
0
10
The French Tech Journal
@index.frenchtechjournal.com.ap.brid.gy
about 2 hours ago
Samsung-led funding gives the French AI champion fresh firepower for frontier research, sovereign infrastructure, and bigger models, addressing doubts that Mistral was falling impossibly behind OpenAI, Anthropic, and China.
Mistral AI Raises €3 Billion at €21 Billion Valuation to Keep Europe in the AI Race

www.frenchtechjournal.com

Mistral AI Raises €3 Billion at €21 Billion Valuation to Keep Europe in the AI Race

Mistral would like to remind everyone that reports of it being "cooked" are premature. The Paris company said Tuesday that it has raised €3 billion at a post-money valuation of more than €21 billion (roughly $24 billion), in what it bills as the largest equity round ever completed by a privately held European technology company. The funding comes three years after its launch. For a startup that began as a small group of ex-DeepMind and ex-Meta researchers with a strong opinion about open models, that is a decent run by any measure. Unless you compare it to private labs in the U.S., in which case it still falls short. Anthropic has raised around $100 billion so far this year ahead of an IPO that could value it at as much as $2 trillion. Reuters pegs OpenAI at $852 billion as it heads to a public listing this year. So, yeah, against that, €3 billion kinda feels like a rounding error. And yet... Seen from inside Europe, the numbers look different. Mistral becomes the continent's second-highest-valued private tech group and its fastest-growing company. The valuation has nearly doubled in a year, from around €12 billion. As for French milestones, the funding is remarkable on several levels. On The French Tech Journal funding tracker, the round puts the ecosystem past €8 billion raised so far in 2026, already topping €7.8 billion for 2025. For vets of the French Tech scene, it was just back in 2017 that €3.1 billion represented the **_total funding for the entire year._** Chart courtesy of Dealroom For the latest round, Samsung Electronics led the funding while the EU-backed Scaleup Europe Fund, managed by EQT, and existing shareholder PSG Equity came in as co-leads. Advent, funds managed by BlackRock and the Grand Duchy of Luxembourg, joined as new investors. Almost everyone already on the cap table re-upped: a16z, ASML, Bpifrance, BNP Paribas CIB, DST Global, General Catalyst, Index Ventures, Lightspeed, Nvidia, Salesforce Ventures. Samsung isn't just a tourist. In an interview with the Financial Times, Mistral AI CEO Arthur Mensch called the investment "another validation" of Mistral's usefulness in high-end manufacturing, following ASML's €1.3 billion cheque last year. Samsung plans to use Mistral's AI to improve its chip manufacturing systems. Although Microsoft agreed in July to spend billions of dollars on Mistral's computing infrastructure in Europe, the Redmond giant apparently did not participate, Mistral AI chief financial officer Johan Bergqvist told Reuters. ## What the money buys Despite a general freakout last month after the announcement that Mistral AI would build compute primarily focused on inference, the company stressed in the funding announcement that: "The round will significantly expand Mistral's **_frontier research_** , which is the foundation underpinning its infrastructure, products and sovereignty." In the FT, Mensch said the company believes it will now have the compute to match the strategy used by Chinese AI labs, which allowed them to match the frontier training of OpenAI and Anthropic at a fraction of the cost. “With this fundraising we will be controlling an amount of compute that is very comparable to what the Chinese labs have,” he said. “We have built Mistral to train models and to scale them. We will continue to do that, and that’s really the purpose of this fundraising.” Mensch told CNBC the money goes toward more infrastructure, including Mistral's own data centers, and that the company will train "bigger and faster models." He said Mistral is on track for $1 billion in annual recurring revenue by year-end, with growth strongest in Asia and North America. Mensch went further with CNBC, saying he expects "to be beating" that figure "if everything happens as they are trending." ## Sovereignty, with an asterisk Mistral AI continues to place sovereignty at the core of its identity and pitch. "The fact that the EU or Europe has to have its own kind of AI provider in the game is important," Bergqvist told Reuters. "We have seen in the past that there is politics involved in the access to these solutions, and it's important that you make sure that you maintain access and do not compromise your kind of supply chain." Mistral's answer is the full stack: open-weight models, the infrastructure they run on, and the products that put them into production. Control means data that stays inside an organization's walls, customizable models, private compute, and auditable systems. The company now works across 20 countries with more than 125 enterprises, including Airbus, ASML, and HSBC. Still, eyebrows were raised in August when Mistral said it would commercialize GLM-5.2, a model built by the Chinese firm Z.ai. Mensch defended it to the FT as giving customers choice while still offering "regional controls" to protect their data. But he also hinted that it may just be a stopgap solution. Mensch told CNBC the models Mistral releases "very soon" will be "very competitive."

0
0
0
🚀 olud.ai
@oludai.bsky.social
about 12 hours ago
🎨 Blind human preference: Qwen-Image-3.0-Pro, the best downloadable image model, is just 43 ELO points behind MAI-Image-2.6 (Microsoft AI). That gap is tiny for open-source. See the ranking on olud.ai. olud.ai/media-leaderboard… #OpenSource #GenerativeAI #StableDiffusion #AI
0
0
2
Techyon news
@techyonai.bsky.social
about 5 hours ago
Running Qwen 3.8 27B at Q4 on 16GB VRAM at 200K CTX at 50t/s A user successfully runs the Qwen 3.8 27B model at Q4 precision on a 16GB RTX 5070Ti, achieving 50 tokens per second with a 200K context window. The s...

techyon.pages.dev

Running Qwen 3.8 27B at Q4 on 16GB VRAM at 200K CTX at 50t/s

A user successfully runs the Qwen 3.8 27B model at Q4 precision on a 16GB RTX 5070Ti, achieving 50 tokens per second with a 200K context window. The setup uses Unsloth's UD-IQ4_XS GGUF quantization an

0
0
0
AlternativeTo
@alternativeto.net
2 days ago
MiniMax launched H3, a multimodal model for creating content using text, image, video, and audio. It generates 2K videos up to 15s with stereo sound. Model weights will be released soon. alternativeto.net/news/20….
A purple background features bold text announcing "MiniMax H3," a next-generation multimodal video model.
0
1
6
Mike Ammerlaan
@mikeamm.bsky.social
about 8 hours ago
Gezel update: mostly just better infrastructure: updates to #llamacpp and #dwarfstar engines for cutting edge model execution, plus better air traffic control of memory. gezel.com/docs/whats-n...
1.26251 — 8 September 2026

gezel.com

1.26251 — 8 September 2026

Engines that reserve memory instead of colliding, llama.cpp v0.4 loading controls, GLM 5.3 on DwarfStar, and catalogs built from a folder of Markdown.

0
0
0
What's New
@whatsnew.fyi
about 12 hours ago
open-design open-design-v0.22.0 🔎 66 PRs · 10 contributors · 7 days — We evaluated over a dozen models using the OpenDesign Harness to compare design quality and cost. DeepSeek V4.1 Flash achieved 98% of top-ranked GPT-6 Astra’s average score at just 1% of its average cost. Explore OpenDesign…
open-design open-design-v0.22.0

whatsnew.fyi

open-design open-design-v0.22.0

🔎 66 PRs · 10 contributors · 7 days — We evaluated over a dozen models using the OpenDesign Harness to compare design quality and cost. DeepSeek V4.1 Flash achieved 98% of top-ranked GPT-6 Astra’s average score at just 1% of its average cost. Explore OpenDesign Arena to find the right model for…

0
0
0
Thomas Wood
@advanced-eschatonics.com
4 days ago
Running glm-5.3-flash at home. Prefills are brutal but MTP is working right out of the box. I'm seriously impressed already.
GLM-5.3-Flash: How to Run Locally | Unsloth Documentation

unsloth.ai

GLM-5.3-Flash: How to Run Locally | Unsloth Documentation

Run the new GLM-5.3-Flash aka ox-alpha model by Z.ai.

1
2
13
Thomas Wood
@advanced-eschatonics.com
4 days ago
okay this model is phenomenal even with IQ1_S. I've still got some room to spare for context so downloading IQ2_XXS. we really have opus-4.8 at home
0
0
7
DEV Community [Unofficial]
@dev.to.web.brid.gy
about 9 hours ago
Leanstral 1.5 is a what..?

dev.to

Leanstral 1.5 is a what..?

I was doing the do out on the internet, and you do, and came across Mistral AI Leanstral 1.5. Which calls itself: > An updated Lean 4 formal proof engineering model optimised for automated theorem proving and autoformalization. 119B total parameters, 6.5B active. I was like, wow, thats a thingymabob of the first order. The last time someone said "formal methods" to me was in 1998, when a bloke called in a computer lab in Surrey. I had not a scooby doo what that was, and here we are more than a quarter of a century later, and solving something cool with formal methods is still on my bucket list. Boo hoo. I had to check the blog to ask what-the-actual-fluffs this is all about Leanstral 1.5: Proof Abundance for All. They are letting you use it for free. So you can go wild having it formally prove you don't have bugs in your mission-critical systems code! Yipee! Well, it turns out I need a formal proof that my Unbounded Viewstamped Replication Revisited code doesn't have many problems. So I thought I'd ask Fable 5.1, GPT 5 and GLM-5.3-Flash to kick the tyres on getting the Leanstral model to write out formal proofs in pretty Greek characters and debug it. Huh you did what!? Well, lets say a model was not good at, I dunno, coding in a rare language, like Felienne Hermans’ Hedy. You could fine-tune a model on that and make a language server that lints the code so that a special model trained on a fast set of examples can happily write code in Arabic and right-to-left with all the different and rare Arabic numerals. That would be so darn cool. Well, Mistal AI have done that for the LeanProver language Lean. They have a VS Code Plugin. On their blog, they show Claude can find real solutions to real problems in that system, but at a very high cost. I told Claude to call Leanstral if it needed help in writing Lean proofs that could be solved. Which it duly did when it got a bit stuck. Here is what it wrote: So there you go! So, like, you know you could, I dunno,

0
0
0
snuow
@snuow.bsky.social
about 14 hours ago
過去に作った動画です。 Qwen3.5:9b を5つのテストで検証!GPT OSS 20B と比較してみた【ローカルLLM】 Alibaba が公開した軽量ローカルLLM「Qwen3.5-9b」を、GPT-OSS 20B・Qwen3 14B と5つのテストで徹底比較... URL: www.youtube.com/watch?v=G…
0
0
0
Simon P. Couch
@simonpcouch.com
6 days ago
We just shipped GLM 5.3 and GLM 5.3 Flash in Posit AI Pass! GLM 5.3 Flash is now the cheapest model available via the subscription (half the price of Gemma 4 26B A4B per-token!) and I've been so, so impressed with it. GLM 5.3 is Ox Alpha, the anonymous model that was popping off on OpenRouter.
Table comparing AI model costs per token, sorted from cheapest to most expensive. Columns: Model, Lab, and Relative Cost Per-Token. GLM 5.3 Flash (Z.ai) is cheapest at 0.05x; Gemma 4 26B (Google, going away soon) at 0.1x; Claude Haiku 4.5 (Anthropic, going away soon) at 0.33x; GLM 5.3 (Z.ai) and GLM 5.2 (Z.ai, going away soon) both at 0.38x; Claude Sonnet 5 (Anthropic) at 0.67x; Claude Sonnet 4.6 (Anthropic) at 1x as the baseline; Kimi K3 (Moonshot AI) at 1x; and Claude Opus (Anthropic) as the most expensive at 1.67x.
1
2
14
Paul Piper
@madppiper.bsky.social
about 18 hours ago
alvins82 benchmarked 10 model/harness combos on the same Three.js scene. Astra 6.0 Max answered fastest (3.6s to first token), then took 37 minutes total to finish. Qwen 3.8 27B on OpenCode did the whole thing in under 9. Fast start and fast finish turned out to be unrelated numbers.
0
0
0
OR13
@or13.io
3 days ago
GLM-5.3 beat Opus 4.8 on Z.ai's own Code Bench: 31.4% vs 29.5%. Ignore that. The number that matters is ~50K output tokens vs ~120K. Agent work is billed in tokens, not accuracy points. Score-per-token is the leaderboard nobody publishes. z.ai/blog/glm-5.3

z.ai

GLM-5.3: Frontier Coding with Emergent Cyber Capabilities

Z.ai's open-weights 753B MoE coding model: 31.4% on Z.ai Code Bench at ~50K output tokens vs Claude Opus 4.8's 29.5% at ~120K.

1
0
4
@aichina.news
about 11 hours ago
Apache 2.0 FLUX LoRA on Huawei’s Ascend hub looks like a low-friction custom-image win — no retraining, no lock-in. But no README, sample output, base model or precision stated. That’s a stub, not a card. Hard to see this beating Qwen-Image’s adapters until Modelers ships eval or a working...
Meet V3OS-nym2: A Lightweight LoRA Adapter for FLUX on Huawei's Ascend Ecosystem

aichina.news

Meet V3OS-nym2: A Lightweight LoRA Adapter for FLUX on Huawei's Ascend Ecosystem

0
0
0