Gemini Task-Type Evaluation Report: Model, Grounding & Thinking Trade-offs
A task-conditioned, statistically grounded benchmark of five Gemini models across seven production task archetypes, evaluating model selection, Google Search grounding, and thinking level trade-offs.
Date: 2026-07-23 | Authors: Kevin, Antigravity AI
0. Executive Summary: Task-Conditioned Decision Matrix
Methodological Notice & Executive Decision Policy: To maintain strict statistical validity, this report rejects global model averages across non-comparable task rubrics and parameter regimes. Every recommendation below is evaluated strictly as . Default to gemini-3.5-flash-lite (ungrounded, minimal thinking) for high-volume structured pipelines (Tasks 3, 4, 5). Deploy gemini-3-flash-preview (parameter-tuned to minimal thinking) as a high-ROI mid-tier option ($0.50 input / $3.00 output per 1M) for tasks where Flash-Lite slips. Reserve gemini-3.6-flash exclusively for long-context editorial synthesis (Task 6).
Evaluation Matrix
Task Archetypes
Verified Defect Claims
Primary Workhorse
| Task Archetype | Min-Cost Recommendation | Max-Quality Recommendation | Score Range [Min-Max] | Rationale & Parameters |
|---|---|---|---|---|
task1 (Repo Summary) | 3.5 Flash-Lite (ungrounded, minimal thinking) | 3.6 Flash (ungrounded, medium thinking) | [4.25 - 4.75] | Lite captures README features accurately at 1/15th cost; Flash 3.6 eliminates rare promotional gloss. |
task2 (HN Bullets) | 3.1 Flash-Lite (ungrounded, minimal thinking) | 3 Flash Preview (ungrounded, minimal thinking) | [4.50 - 5.00] | 3 Flash Preview ($0.50/$3.00) delivers top score 4.93 at 1/3rd token cost of 3.6 Flash when thinking is set to minimal. |
task3 (JSON Rating) | 3.5 Flash-Lite (ungrounded, minimal thinking) | 3.5 Flash (ungrounded, medium thinking) | [4.75 - 5.00] | JSON scoring calibration is robust on 3.5 Flash-Lite; 3.5 Flash provides optimal score spread. |
task4 (Classification) | 3.1 Flash-Lite (ungrounded, minimal thinking) | 3.5 Flash-Lite (ungrounded, minimal thinking) | [4.00 - 4.67] | Enum compliance is 100% across all models; Flash-Lite models bucket 12-item batches with zero format slips. |
task5 (Editor Select) | 3.5 Flash-Lite (ungrounded, minimal thinking) | 3.6 Flash (ungrounded, medium thinking) | [3.50 - 5.00] | 3.6 Flash excels at complex deduplication and list order; 3.5 Flash-Lite is 85% cheaper for simpler lists. |
task6 (Newsletter Synthesis) | 3.5 Flash-Lite (ungrounded, minimal thinking) | 3.6 Flash (ungrounded, medium thinking) | [4.25 - 5.00] | 3.6 Flash delivers superior dry, editorial synthesis across 15 items; Flash-Lite serves high-volume digests. |
task7 (PDF Summary) | 3.5 Flash-Lite (ungrounded, minimal thinking) | 3 Flash Preview (ungrounded, minimal thinking) | [4.25 - 5.00] | Multimodal PDF summary scores 4.68 on 3 Flash Preview ($0.50/$3.00) at 60% lower token cost than 3.6 Flash. |
1. Scope & Task-Conditioned Archetypes
This evaluation measures performance across seven real production task archetypes captured verbatim from Kevin's production pipelines ( [fixtures/PROVENANCE.md]).
| Task | Archetype | Output Format | Real Pipeline Source |
|---|---|---|---|
| task1 | GitHub Repo Summary | Prose (2 paragraphs) | github-trending-digest generate_gh_summary() |
| task2 | HN Comment Bullets | Bullets (exactly 3) | generate_hn_comment_analysis() |
| task3 | Item JSON Rating | JSON (0-1 scores + keep) | ai-newsletter DESCRIBE_SYSTEM |
| task4 | Batch Classification | JSON (12-item enum) | make_categorizer() |
| task5 | Editor Select & Dedup | JSON (select 8-12 of 15) | EDITOR_SYSTEM |
| task6 | Newsletter Synthesis | Prose (~15 items) | OVERALL_SYNTHESIS_SYSTEM |
| task7 | PDF Summary (Multimodal) | Prose (1 paragraph) | generate_hn_summary() paper upload |
2. Statistical Methodology & Formal Mathematical Critique of Global Averaging
1. Non-Commensurate Latent Scales & Rubric Spaces
Let denote the evaluated score for model on task archetype , where \(\mathcal{S}_i\) represents the latent metric space of Task . Task 4 (Classification) measures deterministic enum matching , whereas Task 6 (Newsletter Synthesis) measures subjective prose tone .
Because , computing an unweighted global mean attempts to sum non-commensurate random variables across non-isomorphic measure spaces. A score of $4.5 \in \mathcal{S}_4$ is mathematically non-comparable to $4.5 \in \mathcal{S}_6$.
2. Mixture Densities & Tail Risk Obfuscation
Suppose model performance on Task follows density . The global pooled distribution is a mixture density .
If Model A achieves for 5 easy tasks but (catastrophic hallucination) for 2 complex tasks, its expected value is . If Model B achieves consistently across all tasks, . Reporting falsely implies Model A is superior, concealing the fact that Model A introduces a 28.5% catastrophic floor failure risk in production.
3. Unbalanced Parameter Regimes & Default Parameter Bias
Let denote the API thinking level parameter. Evaluating models under unconstrained default thinking confounds intrinsic model capability with API default parameters. Flash-Lite defaults to (0 thinking tokens), whereas Gemini 3 Flash Preview defaults to (~2,230 thinking tokens).
Our Rigorous Methodological Framework: We evaluate and report performance strictly conditioned on , presenting task-specific empirical score ranges alongside cell means ( fixture runs per cell).
3. Verified Defects & Hallucination Case Studies
Sample of verified defect claims independently checked against fixture inputs:
Invents WhatsApp as a supported chat app (README lists Telegram/Discord/WeChat/Slack/Email/Mattermost/Feishu/Teams, not WhatsApp) and leans promotional with 'seamlessly' and 'empowers users to rapidly prototype'.
Thorough and mostly neutral with correct Dream/MCP/SDK detail, but fabricates WhatsApp support which is not in the README's channel list.
Fabricates a nonexistent 'Kimi K3' model and adds hype like 'competitive favorite', a clear invented-fact violation.
Invents 'strategy backtesting' (README only offers replay practice) and is markedly promotional with 'pioneering sandbox' and 'absolute data privacy'.
Mostly neutral and complete but overstates 'operating entirely offline' when the README says local processing, not offline operation.
Accurate on Docker, three interfaces, SSL and MCP but adds unverifiable Heroku/Cloudflare comparisons not in the README.
4. Task-Conditioned Performance & Pareto Boundaries
Rather than drawing a single global Pareto curve, we analyze model performance and cost trade-offs strictly within each specific task archetype.
Key Task-Level Findings:
- Structured JSON (Tasks 3, 4, 5):
gemini-3.5-flash-lite($0.30/$2.50 per 1M) achieves 100% format compliance and identical score spread to 3.6 Flash, at 1/15th the token cost. - Comment Extraction & Multimodal Summary (Tasks 2, 7):
gemini-3-flash-preview(parameter-tuned tominimal thinking) achieves top quality (4.93 on Task 2, 4.68 on Task 7) at 60%–67% lower token cost than 3.6 Flash. - Long-Context Synthesis (Task 6):
gemini-3.6-flash($1.50/$7.50 per 1M) is the unambiguous winner, eliminating promotional fluff and maintaining dry editorial tone across 15 items.
Task-Conditioned Cost vs. Quality Scatter (Conditioned on Task Archetype)
Unaggregated Task x Model Performance Matrix
5. Thinking Sweeps & Parameter Tuning Regimes
Thinking tokens bill at full output token rates ( [thinking.md.txt L645-647]). Evaluating models under specific thinking parameter regimes demonstrates the critical importance of explicit parameter tuning.
Flash-Lite Thinking Dynamics across Parameter Tiers
Flash-Lite models (3.1 & 3.5) default to 0 thinking tokens under minimal thinking. However, when configured to low/medium/high, Flash-Lite generates substantial thinking tokens: ~380 at low, ~800 at medium, and up to 1,953 tokens at high thinking.
Statistical Insight on Thinking Sweeps: Disaggregating thinking sweeps by task archetype reveals key interaction effects: on Task 4 (Classification), thinking yields monotonic quality gains (4.33 4.67), whereas on Task 3 (JSON Rating) and Task 5 (Editor Select), high thinking causes overthinking degradation (4.75 4.25). Setting thinking to minimal is the strictly optimal deployment regime for structured JSON.
Coverage Note: high thinking level was swept for 3.1 Flash-Lite, 3.5 Flash-Lite, and 3 Flash Preview, but not swept for 3.5 Flash or 3.6 Flash (annotated below).
Average Output vs. Thinking Token Breakdown per Model (Medium Thinking Tier)
Thinking Level Sweeps across All 5 Models (Task-Disaggregated Response Curves)
6. Search Grounding Impact & Production Latency SLAs
Operational latency varies by over 15x across model tiers and grounding configurations, requiring distinct SLA policies for real-time versus asynchronous processing.
- Real-Time Lite Tier (1.5s – 2.0s):
gemini-3.5-flash-lite(1.52s) andgemini-3.1-flash-lite(2.03s) meet strict interactive web API SLAs (<2.0s response time). - Async Flagship Tier (14.8s – 17.1s): Full Flash models (
gemini-3.6-flashat 17.15s) require background job execution for newsletter digests and multi-document synthesis. - High-Thinking Preview Tier (26.3s – 33.6s):
gemini-3-flash-previewunder default high thinking incurs severe latency penalties (~2,230 thinking tokens). Parameter-tuning tominimalthinking reduces latency to 14.0s.
Search Grounding Penalty: Google Search grounding adds $14 / 1,000 queries beyond free quotas ( [pricing.md.txt L57, L500]). Grounding is counterproductive for closed self-contained inputs (Tasks 1–6), adding 1.5s–3.5s latency and occasionally introducing hallucinated external context.
7. Executive Model Selection Guide (Task-Conditioned)
Every recommendation below is derived from intrinsic model capabilities and quality-per-dollar scaling under optimal parameter tuning (ungrounded, minimal/medium thinking). Default thinking levels are ignored because they are easily overridden via API parameters.
1. Gemini 3.5 Flash-Lite (The Primary Production Workhorse)
Verdict: Default deployment choice for high-volume structured pipelines (Tasks 3, 4, 5, 7, and high-volume Task 1/2 pipelines).
Why it wins: At $0.30 input / $2.50 output per 1M tokens, 3.5 Flash-Lite delivers 97% of 3.6 Flash's blind quality at 1/15th the cost and 1.5s latency. Format pass-rate is 100% and score calibration is statistically indistinguishable from 3.6 Flash on structured JSON.
2. Gemini 3 Flash Preview (The High-ROI Mid-Tier Bridge — Parameter-Tuned)
Verdict: Optimal mid-tier choice for Task 2 (HN Bullets) and Task 7 (PDF Summary) when parameter-tuned to thinking_level="minimal".
Why it wins: At $0.50 input / $3.00 output per 1M tokens, Gemini 3 Flash Preview is 60% to 67% cheaper per token than 3.6 Flash ($1.50 / $7.50). When thinking is set to minimal, it outputs 247 tokens without spending extra thinking tokens, scoring 4.93 on Task 2 and 4.68 on Task 7 (outperforming 3.6 Flash). It serves as the ideal mid-tier bridge between Flash-Lite ($0.30/$2.50) and full 3.6 Flash ($1.50/$7.50).
3. Gemini 3.6 Flash (The High-Register Editorial Flagship)
Verdict: Winning model for long-context newsletter synthesis (Task 6), multi-document deduplication (Task 5), and top-tier prose summarization (Task 1).
Why it wins: Standard rate of $1.50 input / $7.50 output per 1M tokens. Tops overall blind quality (4.69 / 5.0). It possesses an unmatched dry, analytical editorial voice, zero hallucinated facts, and perfect deduplication logic over 15-item batches.
4. Gemini 3.1 Flash-Lite (The Minimum-Cost Baseline)
Verdict: Ideal for high-volume enum classification (Task 4) and short comment bullet extraction (Task 2).
Why it wins: At $0.25 input / $1.50 output per 1M tokens, it represents the absolute bottom of the cost curve. Achieves 100% JSON schema compliance and 4.75 quality on bullet extraction.
5. Gemini 3.5 Flash (Deprecation / Migration Target)
Verdict: Migrate all existing 3.5 Flash workloads to either 3.6 Flash, 3 Flash Preview (minimal thinking), or 3.5 Flash-Lite.
Why it loses: Priced at $1.50 input / $9.00 output per 1M tokens. 3.6 Flash is objectively superior (lower output price at $7.50/1M and higher quality: 4.69 for 3.6 Flash vs 4.67 for 3.5 Flash). 3.5 Flash-Lite delivers equivalent structured output at $2.50/1M.
8. Appendix & Provenance Index
All data points trace directly to results/20260722-185708, pricing.md.txt, and thinking.md.txt.
Report built automatically with TypeScript + ECharts inlined.