Merlin Research

non-profit
Activity Feed

AI & ML interests

Independent AI safety lab. Stockholm, Sweden. We test deployed LLM agents under adversarial conditions and measure behavioral alignment in production โ€” not in controlled benchmarks.

Recent Activity

squ11z1ย  updated a model about 2 months ago
Merlin-Research/Qwen3.5-4B-Safety-Thinking
squ11z1ย  updated a model about 2 months ago
Merlin-Research/Pluto
squ11z1ย  updated a model about 2 months ago
Merlin-Research/HybridIntelligence-0.5B
View all activity

DedeProGamesย 
posted an update about 15 hours ago
view post
Post
697
๐Ÿš€ OxCoder-9B โ€” a lightweight agentic coding model, now on HF!

Introducing OxCoder-9B, a 9B parameter model built for long-horizon tasks, agentic coding, and agentic reasoning. Despite its compact size, it delivers frontier-level performance in Agentic Terminal and Agentic Coding, rivaling models many times its size.

Highlights:
- Trained on frontier agent traces โ€” distilled from Fable-5.1 and GLM-5.3 agentic coding trajectories across Claude Code, OpenCode, and Codex
- 262K native context โ€” handles complex, multi-file codebases and long-horizon reasoning tasks with ease
- Error recovery โ€” learns read-before-write patterns, responds to LSP diagnostics, and applies minimal edit diffs instead of full rewrites
- Strong front-end reasoning โ€” deep understanding of UI logic, component architecture, and web-native patterns, rare in sub-10B models

Benchmarks (vs. Ornith-1.5-9B, Ornith-1.0-9B, Qwen3.5-9B, and Gemma-4-31B):
- Terminal-Bench 2.1 (Terminus-2): 49.6
- Terminal-Bench 2.1 (Claude Code): 50.8
- SWE-bench Verified: 73.5
- SWE-bench Pro: 49.1
- NL2Repo: 36.2
- HLE (no tools): 21.2
- HLE (with tools): 32.8
- GPQA Diamond: 86.9
- MCP-Atlas: 56.7
- BrowseComp: 57.4
- ClawEval: 67.8

All OxCoder-9B results are averaged over five independent runs. Built on Qwen/Qwen3.5-9B, released under Apache 2.0.

๐Ÿ”— OrionLLM/OxCoder-9B
  • 1 reply
ยท
DedeProGamesย 
posted an update 6 days ago
view post
Post
115
Possible Kiyo sizes

Kiyo-230M (Kiyo-Ultra)
Kiyo-135M (Kiyo-Plus)
Kiyo-65M (Kiyo-Go)
Kiyo-15M (Kiyo-Air)
Kiyo-2M (Kiyo-Pico)
DedeProGamesย 
posted an update about 1 month ago
view post
Post
2461
DynamicMind-Mini-Instruct just released! A ~8.9m parameters model instruction-tuned based on DynamicMind-Mini
  • 6 replies
ยท
DedeProGamesย 
posted an update about 1 month ago
view post
Post
1561
๐Ÿš€ Introducing the GRM-3.2 Family

The GRM-3.2 family is a new generation of reasoning-focused models from OrionLLM, purpose-built for long-horizon agentic tasks, extremely difficult reasoning problems, advanced coding, and local AI workflows across a wide range of hardware constraints.

GRM-3.2-Sky is the flagship model in the family: a 35B-A3B Mixture-of-Experts model built on the Ornith-1.0-35B architecture, designed for elite structured reasoning, complex multi-file coding, advanced mathematics, and sustained coherence across extended agentic workflows. It represents a substantial leap in long-horizon task capability over its predecessor, GRM-2.6-Plus.

GRM-3.2-Cliff is the mid-sized workhorse: a 9B-parameter model optimized for long-horizon agentic tasks and difficult reasoning in low-to-mid GPU environments. It delivers strong multi-step planning, debugging, and terminal-agent performance without demanding flagship-level hardware.

GRM-3.2-Turf is the lightweight edge model: a 1.2B-parameter model based on the LiquidAI/LFM2.5-1.2B-Thinking architecture, engineered for efficient on-device execution, high-fidelity instruction following, and robust tool use on mobile, embedded, and other resource-constrained hardware.

All three models are designed for users who need dependable reasoning engines that can maintain goal-directed behavior, planning quality, and task fidelity across many stepsโ€”whether on a server, a local workstation, or an edge device.

Models:
GRM-3.2-Sky: OrionLLM/GRM-3.2-Sky
GRM-3.2-Cliff: OrionLLM/GRM-3.2-Cliff
GRM-3.2-Turf: OrionLLM/GRM-3.2-Turf

Organization:
OrionLLM

  • 2 replies
ยท
DedeProGamesย 
posted an update about 1 month ago
view post
Post
1058
๐Ÿš€ Introducing the GRM-3.2 Family

The GRM-3.2 family is a new generation of reasoning-focused models from OrionLLM, purpose-built for long-horizon agentic tasks, extremely difficult reasoning problems, advanced coding, and local AI workflows across a wide range of hardware constraints.

GRM-3.2-Sky is the flagship model in the family: a 35B-A3B Mixture-of-Experts model built on the Ornith-1.0-35B architecture, designed for elite structured reasoning, complex multi-file coding, advanced mathematics, and sustained coherence across extended agentic workflows. It represents a substantial leap in long-horizon task capability over its predecessor, GRM-2.6-Plus.

GRM-3.2-Cliff is the mid-sized workhorse: a 9B-parameter model optimized for long-horizon agentic tasks and difficult reasoning in low-to-mid GPU environments. It delivers strong multi-step planning, debugging, and terminal-agent performance without demanding flagship-level hardware.

GRM-3.2-Turf is the lightweight edge model: a 1.2B-parameter model based on the LiquidAI/LFM2.5-1.2B-Thinking architecture, engineered for efficient on-device execution, high-fidelity instruction following, and robust tool use on mobile, embedded, and other resource-constrained hardware.

All three models are designed for users who need dependable reasoning engines that can maintain goal-directed behavior, planning quality, and task fidelity across many stepsโ€”whether on a server, a local workstation, or an edge device.

Models:
GRM-3.2-Sky: OrionLLM/GRM-3.2-Sky
GRM-3.2-Cliff: OrionLLM/GRM-3.2-Cliff
GRM-3.2-Turf: OrionLLM/GRM-3.2-Turf

Organization:
OrionLLM

DedeProGamesย 
posted an update about 2 months ago
view post
Post
230
who want GRM-2.6 with 9b?
DedeProGamesย 
posted an update 3 months ago
DedeProGamesย 
posted an update 4 months ago
view post
Post
378
๐Ÿš€ Introducing the GRM-2.6 Family

The GRM-2.6 family is a new generation of reasoning-focused models from Orion LLM Labs, built for difficult tasks, coding, STEM, terminal agents, and advanced local AI workflows.

GRM-2.6-Plus is the main high-capability model in the family: a 27B-class reasoning model based on Qwen3.6, designed for strong structured reasoning, coding, agentic use, and practical local deployment.

GRM-2.6-Opus builds on GRM-2.6-Plus as a merge with an Opus-style reasoning distilled model, improving structured reasoning behavior, terminal-agent workflows, coding ability, and complex problem solving.

Both models are designed for users who want powerful reasoning models that remain practical for research, local inference, coding, and agent experiments.

Models:
GRM-2.6-Plus: OrionLLM/GRM-2.6-Plus
GRM-2.6-Opus: OrionLLM/GRM-2.6-Opus

Organization:
OrionLLM
DedeProGamesย 
posted an update 4 months ago
view post
Post
5191
๐Ÿš€ Introducing the GRM-2.6 Family

The GRM-2.6 family is a new generation of reasoning-focused models from Orion LLM Labs, built for difficult tasks, coding, STEM, terminal agents, and advanced local AI workflows.

GRM-2.6-Plus is the main high-capability model in the family: a 27B-class reasoning model based on Qwen3.6, designed for strong structured reasoning, coding, agentic use, and practical local deployment.

GRM-2.6-Opus builds on GRM-2.6-Plus as a merge with an Opus-style reasoning distilled model, improving structured reasoning behavior, terminal-agent workflows, coding ability, and complex problem solving.

Both models are designed for users who want powerful reasoning models that remain practical for research, local inference, coding, and agent experiments.

Models:
GRM-2.6-Plus: OrionLLM/GRM-2.6-Plus
GRM-2.6-Opus: OrionLLM/GRM-2.6-Opus

Organization:
OrionLLM
  • 2 replies
ยท
DedeProGamesย 
posted an update 4 months ago
view post
Post
8595
GRaPE 2 Pro is now available.

SL-AI/GRaPE-2-Pro

This is the flagship model of the GRaPE 2 family and the largest model I have trained to date, sitting at 27B parameters. It is built on Qwen3.5-27B and trained on a closed-source proprietary dataset, with roughly half of post-training focused on code and the rest split between STEAM subjects and structured logical reasoning. It punches seriously above its weight class.

GRaPE 2 Pro supports multimodal input (image + text) and features 6 thinking modes via the <thinking_mode> tag. This gives you real control over how hard the model thinks, from skipping the reasoning phase entirely with minimal, all the way up to xtra-Hi for deep, extended thought on hard problems. For most agentic use, auto or low is the move to keep things snappy.

It also runs on consumer hardware. You can get it going with as low as 12GB of VRAM on a quantized build.

If you want to try it out and give feedback, that would be really appreciated. Email us at contact@skinnertopia.com
  • 1 reply
ยท
DedeProGamesย 
posted an update 5 months ago
view post
Post
3489
๐Ÿ”ฅ GRM-2.5 - The most POWERFUL model for local inference

The GRM-2.5 is the newest model from Orion LLM Labs. It has consistent RAW reasoning and is capable of generating very precise responses, similar to large models, while maintaining a parameter size of 4b.

The GRM-2.5 family consists of these models:
OrionLLM/GRM-2.5 (4b)
OrionLLM/GRM-2.5-Air (0.8b)

Furthermore, the GRM-2.5 is the best option for local agentic environments, being very good in code, terminal agent, etc. It is capable of generating 1000 lines of consistent code and programming like large models.
The GRM-2.5 is the best base for FineTune to date and has vision, which means it can interpret images and videos.
  • 1 reply
ยท
DedeProGamesย 
posted an update 5 months ago
view post
Post
3062
๐Ÿ”ฅ GRM2 - The small one that surpasses the big ones.
What if a 3-parameter model can beat a 32-parameter model in every benchmark? We prove that it can.
GRM2 is a 3b params model based on the llama architecture, trained for long reasoning and high performance in complex tasks - the first 3b params model to outperform qwen3-32b in ALL benchmarks, and outperform o3-mini in almost all benchmarks.
๐Ÿค— Model: OrionLLM/GRM2-3b
The first 3b params model to generate over 1000 lines of code and achieve a score of 39.0 in xBench-DeepSearch-2510.

๐Ÿš€ Chat with GRM:
https://huggingface.co/spaces/DedeProGames/GRM2-Chat

๐Ÿ† Download official GGUFs: OrionLLM/GRM2-3b-GGUF