$ grep -r "evaluation" ./posts/

# evaluation

Все ai claude-code llm open-source agents anthropic productivity tips developer-tools claude coding mcp openai tools cursor gemini google api
security codex cli automation ai-agents workflow testing pricing voice models ide qwen comparison coding-tools ai-tools skills tokens multimodal openrouter ai-models plugins gpt ai-coding xai grok cybersecurity leak alibaba benchmarks tdd coding-agent playwright orchestration codex-cli multi-agent context-window coding-agents memory python chatgpt stealth-models google-io-2026 gpt-5-6 moe research git openclaw ralph-loop autonomous-coding github ios swift xcode computer-use gpt-5.4 code-review browser-automation unity game-development context-engineering vibe-coding web-scraping browser gemma china tts glm hunter-alpha deepseek video-generation owl-alpha protocol nvidia vision ollama gpt-5.6 local fable prompt-engineering coding-assistant benchmark devtools deep-research terminal qa php laravel assistant worktrees docker parallel-development oauth websocket context-management mobile copilot perplexity multi-model image-generation remotion video shorts instagram tiktok permissions future code-intelligence knowledge-graph future-of-programming opinion hooks xctest commands local-ai liquid-ai privacy fast-mode copilot-cli macos linux windows machine-learning cron scheduled-tasks effort settings godot unreal-engine search-api tavily exa agent-teams opus-4.6 expo cowork remote-control plugin google-colab responsive-design frontend telegram discord channels astral superapp kimi licensing documentation prompts figma design web-development demo gamedev gemini-cli speech scraping self-improvement ultraplan debugging function-calling free-tools elevenlabs infrastructure configuration skill microsoft dotnet cost-optimization nous-research gpt-6 llama healer-alpha elephant-alpha gpt-5-5 tmux stealth-launch fal elixir linear rust tencent voice-cloning reasoning nemotron mythos policy dense-model game-dev open-beta sonnet gpt-55 spacex managed-agents realtime subq subquadratic long-context transformers finance edge-ai rag vector-search notion typescript workers malware chrome leaks veo lmarena fingerprinting api-pricing onboarding opus-4-8 robotics world-models physical-ai minimax free-models ocr baidu document-ai release-tracker gemini-35-pro ml amazon data-labeling amd hardware local-llm llama-cpp code-quality interpretability ai-safety meta apple lawsuit curl writing pentesting career safety wordpress vulnerability evaluation
coding-benchmarks-retrieval.md
80% на SWE-bench это не 80% решённых задач: как кодинг-лидерборды меряют ретривал, а не кодинг
> · 9 мин

80% на SWE-bench это не 80% решённых задач: как кодинг-лидерборды меряют ретривал, а не кодинг

Четыре работы за месяц показали, что кодинг-лидерборды меряют ретривал и обвязку, а не кодинг: OpenAI зарубила два своих бенчмарка, Cursor нашёл 63% найденных фиксов, RuBench поймал GPT-5.6 на добывании ответов с диска. Плюс инструкция, как честно померить агента на своём репозитории.

ai llm benchmarks coding