I Tested 5 Strong LLMs on Long-Running Prompts
7-minute video for Premium members. Experiment: what if I benchmark models not on ONE prompt, but on the set of FOUR prompts? I tried this approach with 5 top LLMs, which scored the best?
Fresh insights on AI-powered development
7-minute video for Premium members. Experiment: what if I benchmark models not on ONE prompt, but on the set of FOUR prompts? I tried this approach with 5 top LLMs, which scored the best?
After looking my LLM benchmarks table, many people ask how exactly I am testing models, and what are the evaluation criteria. To not repeat it in every video, I decided to publish this article, explaining my metholodogy.
12-minute video for Premium members. Luna is much cheaper than Sol, so can we save money/time to prepare the plan with Sol, give implementation to Luna, and then review with Sol again? Let me show you the numbers from my experiment.
I've been using Grok 4.5 in OpenCode for my LLM benchmarking, but now, as Cursor is bought by the same company, would it make sense to use Cursor with Grok? I've tried, and here are the summarized results, in the video.
6-minute video for Premium members. I decided to test whether harness can change the results - so is DeepSeek v4 Flash in Codex CLI better/worse in terms of quality/price/speed than in OpenCode?
People say local LLMs are getting better for coding. I decided to test how they work in practice, with a few experiments, without investing in extra hardware. In this article, I will show tests on `google/gemma-4-e4b` and `Qwen2.5-Coder-14B-Instruct-GGUF` models.
8-minute video for Premium members. I added Kimi K3 to the previously-executed benchmark, where frontier LLMs had to fix 3 complex bugs in big open-source projects.
16-minute video for Premium members. I gave the frontier models the tasks to not generate new code, but to fix existing code in the very old open-source projects, fixing real issues reported by their users.
11-minute video for Premium members. If you want to protect from AI agent running "rm -rf" or "drop table" as a part of the prompt, this tool will help you.
17-minute video for Premium members. I tested Ponytail skill in Claude Code on three different projects and prompts, comparing the speed, token usage, and the code quality.