The request
Original + machine translationLive solution comparison
Only the selected request runs. Results are reused for the same code version. Each solution starts from independent seeded memory; earlier requests' writes are not replayed. Scores are individual Bench 12 local grades, not full benchmark composites. Tool cases can receive full credit for correct tool behavior even when no answer is found. The official runner sends wire bench_version 9 (the saved probe shows 12) and a live mock-tool endpoint.
Checking runtime version...
Green cards with the same MATCH number identify the same memory: matching user, ID and English original. Each solution is matched against the generator context independently. Other records keep their default background.