fix: align collect_frontier token budget with driver (8k)
kimi-k3 spends heavily on hidden reasoning tokens that draw from the same budget as the visible answer. 3000-token budget resulted in empty responses on hard prompts. Increased to 8000 to accommodate reasoning overhead while ensuring room for substantive visible answers. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -64,11 +64,11 @@ def collect_frontier(items: list[dict], out: Path, model: str = "moonshotai/kimi
|
||||
for it in items:
|
||||
if (it["id"], system) in done:
|
||||
continue
|
||||
# max_tokens=3000, not 1000: kimi-k3 emits hidden reasoning tokens that
|
||||
# draw from the same budget as the visible answer even with
|
||||
# reasoning.enabled=False; 1000 was observed to exhaust on harder
|
||||
# prompts and return empty content (finish_reason="length").
|
||||
text, _ = chat(client, model, [{"role": "user", "content": it["prompt"]}], max_tokens=3000)
|
||||
# max_tokens=8000: kimi-k3 spends heavily on hidden reasoning tokens; 8000
|
||||
# leaves room for the visible answer — 3000 produced empty responses on hard
|
||||
# prompts (finish_reason="length"). Hidden reasoning tokens draw from the
|
||||
# same budget as the visible answer even with reasoning.enabled=False.
|
||||
text, _ = chat(client, model, [{"role": "user", "content": it["prompt"]}], max_tokens=8000)
|
||||
if not text or not text.strip():
|
||||
continue # leave undone for a future rerun rather than write empty content
|
||||
f.write(json.dumps({"prompt_id": it["id"], "system": system, "response": text}) + "\n")
|
||||
|
||||
Reference in New Issue
Block a user