research.ibm.com
How llm-d makes the most of the hardware you already have
IBM Research and Red Hat deployed a 753B open model on H100 GPUs, serving thousands of concurrent coding agents at 5-10x lower cost than commercial APIs.
Tracks discussion of the current best open-weight/open-source LLM: new releases, benchmarks, leaderboards, and head-to-head comparisons across DeepSeek, Qwen, Llama, Mistral, GLM, Kimi, MiniMax, GPT-OSS, Gemma and other open models.
Feed on Blueskyresearch.ibm.com
How llm-d makes the most of the hardware you already have
IBM Research and Red Hat deployed a 753B open model on H100 GPUs, serving thousands of concurrent coding agents at 5-10x lower cost than commercial APIs.
igeek.gamer-geek-news.com
Original post on igeek.gamer-geek-news.com
techyon.pages.dev
Running Qwen 3.8 27B at Q4 on 16GB VRAM at 200K CTX at 50t/s
A user successfully runs the Qwen 3.8 27B model at Q4 precision on a 16GB RTX 5070Ti, achieving 50 tokens per second with a 200K context window. The setup uses Unsloth's UD-IQ4_XS GGUF quantization an
arxiv.org
Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe
Data-sovereignty regulations increasingly require public institutions to deploy open-source, on-premise LLM agents that chain multiple tool-calls across live government APIs. However, open-source models consistently underperform in this multi-step setting, and no existing benchmark measures the gap…
gezel.com
1.26251 — 8 September 2026
Engines that reserve memory instead of colliding, llama.cpp v0.4 loading controls, GLM 5.3 on DwarfStar, and catalogs built from a folder of Markdown.
www.tomshardware.com
Benchmarking Qwen 3.8 27B on RTX 5090 and beyond — VRAM capacity alone can't overcome severe software and inference engine bottlenecks
Following the release of Qwen 3.8 27B, we put our trusty hardware to the test to see which hardware might be best suited for running this open-weight AI model.
whatsnew.fyi
open-design open-design-v0.22.0
🔎 66 PRs · 10 contributors · 7 days — We evaluated over a dozen models using the OpenDesign Harness to compare design quality and cost. DeepSeek V4.1 Flash achieved 98% of top-ranked GPT-6 Astra’s average score at just 1% of its average cost. Explore OpenDesign Arena to find the right model for…
unsloth.ai
GLM-5.3-Flash: How to Run Locally | Unsloth Documentation
Run the new GLM-5.3-Flash aka ox-alpha model by Z.ai.
z.ai
GLM-5.3: Frontier Coding with Emergent Cyber Capabilities
Z.ai's open-weights 753B MoE coding model: 31.4% on Z.ai Code Bench at ~50K output tokens vs Claude Opus 4.8's 29.5% at ~120K.
aichina.news
Meet V3OS-nym2: A Lightweight LoRA Adapter for FLUX on Huawei's Ascend Ecosystem