Fetching from the wire…
Infra2026-08-15 · source-backed
The worked example trains Amazon Nova Lite 2.0 on 500 programming tasks using GRPO with LoRA on SageMaker HyperPod, with a four-component reward: correctness at 1.0, asking-before-coding at 0.6, a guessing penalty at 0.4, loop detection at 0.2. AWS The most useful line is diagnostic: "a component with near-zero within-group variance contributes nothing to learning." A reward term can look healthy in aggregate metrics while teaching the model nothing.
Each link below shares sources, entities, or timing with this story.
OpenAI partners with AWS / Shared entity: AWS / Same source domain / Earlier coverage / Tension
Linked by a graph relationship (OpenAI partners with AWS); both cover AWS; reported by the same outlet (aws.amazon.com).
Bedrock built by AWS / Shared entity: AWS / Same source domain / Earlier coverage
Linked by a graph relationship (Bedrock built by AWS); both cover AWS; reported by the same outlet (aws.amazon.com).
Linked by a graph relationship (Bedrock built by AWS); both cover AWS; reported by the same outlet (aws.amazon.com).
Linked by a graph relationship (Bedrock built by AWS); both cover AWS; reported by the same outlet (aws.amazon.com).
Linked by a graph relationship (Bedrock built by AWS); both cover AWS; reported by the same outlet (aws.amazon.com).
Claude Code supports AWS / Shared entity: AWS / Earlier coverage / Tension
Linked by a graph relationship (Claude Code supports AWS); both cover AWS; earlier AWS coverage from 2026-08-13.
AWS partners with Cisco / Shared entity: AWS / Earlier coverage / Tension
Linked by a graph relationship (AWS partners with Cisco); both cover AWS; earlier AWS coverage from 2026-06-27.
Amazon released AWS / Shared entity: AWS / Earlier coverage / Tension
Linked by a graph relationship (Amazon released AWS); both cover AWS; earlier AWS coverage from 2026-04-30.