Latest Articles

Fresh insights on AI-powered development

PREMIUM
Sol vs Opus: Which One Reviews Code Better?
Aug 29, 2026 • 7 min video

Sol vs Opus: Which One Reviews Code Better?

7-minute video for Premium members. I tested Sol vs Opus as code REVIEWERS, running the same prompt 8 times: 4 on Opus, 4 on Sol, also on different thinking levels of Medium/High. Is there a clear winner?

PREMIUM
I Tested 5 Strong LLMs on Long-Running Prompts
Aug 24, 2026 • 7 min video

I Tested 5 Strong LLMs on Long-Running Prompts

7-minute video for Premium members. Experiment: what if I benchmark models not on ONE prompt, but on the set of FOUR prompts? I tried this approach with 5 top LLMs, which scored the best?

LLM Coding Leaderboard: My Methodology and Scoring Formulas
Aug 21, 2026 • 12 min read

LLM Coding Leaderboard: My Methodology and Scoring Formulas

After looking my LLM benchmarks table, many people ask how exactly I am testing models, and what are the evaluation criteria. To not repeat it in every video, I decided to publish this article, explaining my metholodogy.

PREMIUM
Save Money: GPT-5.6-Luna with Sol Review (Experiment)
Aug 17, 2026 • 12 min video

Save Money: GPT-5.6-Luna with Sol Review (Experiment)

12-minute video for Premium members. Luna is much cheaper than Sol, so can we save money/time to prepare the plan with Sol, give implementation to Luna, and then review with Sol again? Let me show you the numbers from my experiment.

I Tried Using Grok 4.5 in OpenCode vs Cursor
Aug 10, 2026 • 4 min video

I Tried Using Grok 4.5 in OpenCode vs Cursor

I've been using Grok 4.5 in OpenCode for my LLM benchmarking, but now, as Cursor is bought by the same company, would it make sense to use Cursor with Grok? I've tried, and here are the summarized results, in the video.

PREMIUM
I Tried Using Deepseek-v4-Flash in Codex CLI vs OpenCode
Aug 10, 2026 • 6 min video

I Tried Using Deepseek-v4-Flash in Codex CLI vs OpenCode

6-minute video for Premium members. I decided to test whether harness can change the results - so is DeepSeek v4 Flash in Codex CLI better/worse in terms of quality/price/speed than in OpenCode?

I Tried Two Local LLMs for Coding on my MacBook Pro M4
Aug 2, 2026 • 5 min read

I Tried Two Local LLMs for Coding on my MacBook Pro M4

People say local LLMs are getting better for coding. I decided to test how they work in practice, with a few experiments, without investing in extra hardware. In this article, I will show tests on `google/gemma-4-e4b` and `Qwen2.5-Coder-14B-Instruct-GGUF` models.

PREMIUM
Frontier LLM Benchmark: Fable vs Sol vs Opus - on 3 Projects
Jul 17, 2026 • 16 min video

Frontier LLM Benchmark: Fable vs Sol vs Opus - on 3 Projects

16-minute video for Premium members. I gave the frontier models the tasks to not generate new code, but to fix existing code in the very old open-source projects, fixing real issues reported by their users.

Povilas Korop

Get Weekly AI Coding News

Sent every Wednesday. No spam, ever. Unsubscribe anytime.