Fetching from the wire…
Public story · 2026-09-18 · high
He labeled 4,290 Reddit comments once with Gemini, then fine-tuned a 459M-parameter model that runs locally with no per-call cost.
Why now: The comparison matters now because per-call extraction pricing is still the default most teams haven't questioned.
Peter Vijeh spent $11.50 total to build a Reddit knife-brand extractor, replacing per-call API pricing, according to his GLiNER fine-tuning write-up. The split was $9 for Gemini labeling and $2.50 for GPU time. After that, the model runs locally with no per-call cost, no matter the volume.
He had Gemini 3.1 Pro label 4,290 Reddit comments once, tagging three attributes: knife brands, models, and steel types. Character offsets came from code, not manual annotation, so the labeling pass covered the full dataset without per-example review.
That labeled set trained GLiNER large v2.5, a 459M-parameter model, on a single Tesla T4 GPU. The fine-tune scored 0.83 F1 on held-out validation, with brand extraction at 0.904 F1 and material recall at 0.911.
The catch is the schema has to hold still. Three fixed attributes made one labeling pass enough to cover the space. A schema that keeps changing means relabeling and retraining every time, and the economics stop working.
Each link below shares sources, entities, or timing with this story.
A June 18 model-tracking roundup reports Google set Gemini 2.5 Flash as the default across its consumer Gemini products, prioritizing latency and cost. This is single-source as of writing, so flag it pending the official Gemini blog, but it lines up with Google's other June 18...
1. Defend Against Vibeware DDoD — Behavioral process monitoring for AI-generated polyglot malware. Monitor Living Off Trusted Services patterns. Bitdefender 2. OS-Level Agent Sandbox — Kernel-level sandboxing (Seatbelt/Bubblewrap/AppContainer) for agentic workflows. Applicatio...
msitarzewski/agency-agents added 446 stars today, packaging personas across "divisions" (frontend specialists, community experts, fact-checkers, reality checkers), each defined with a voice, a process, and concrete deliverables rather than a generic prompt template (GitHub). I...
Luu's September 1 post drew 852 points and over 1,000 comments, walking predictions from February 2024 through November 2025 after removing unfalsifiable and tautological ones. Misses include "AI has peaked" (Feb 2024), Meta "dying" (Nov 2024) against revenue going $135B to $2...
Google launched agentic video understanding across Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite (Google). Instead of scanning a video start to finish at a fixed sample rate, the model runs an internal loop deciding what to watch, at what speed, and through which channel: fra...
Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image) generates images in as little as four seconds and is live in AI Studio, the Gemini API, AI Mode, and the Gemini app. Gemini Omni Flash entered public preview for video generation and conversational editing at $0.10 per second, m...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.