{"researchdata_availability":"no","main_file_link":[{"open_access":"1","url":"https://doi.org/10.52202/085713-3163"}],"author":[{"first_name":"Roman","last_name":"Garipov","full_name":"Garipov, Roman"},{"first_name":"Fedor","last_name":"Velikonivtsev","full_name":"Velikonivtsev, Fedor"},{"first_name":"Ivan","last_name":"Ermakov","full_name":"Ermakov, Ivan"},{"first_name":"Ruslan","full_name":"Svirschevski, Ruslan","last_name":"Svirschevski"},{"first_name":"Vage","full_name":"Egiazarian, Vage","last_name":"Egiazarian","id":"77451e76-92b2-11ef-a4d1-8dbaa06e16ad"},{"first_name":"Max","last_name":"Ryabinin","full_name":"Ryabinin, Max"}],"OA_type":"free access","supplementarymaterial":"yes","date_created":"2026-09-06T22:02:00Z","OA_place":"publisher","fulldoi":"https://doi.org/10.52202/085713-3163","acknowledgement":"We would like to express our sincere gratitude to Denis Mazur for his valuable contributions to the\r\nimplementation of API calls used in Algorithm 1 and for supporting inference with the Llama 405B\r\nmodel. We are also thankful for his positive influence on the overall atmosphere and team morale\r\nthroughout the course of this project.","das_tickbox":"0","publication_identifier":{"isbn":["9798331338275"],"issn":["1049-5258"]},"abstract":[{"lang":"eng","text":"We introduce AutoJudge, a method that accelerates large language model (LLM) inference with task-specific lossy speculative decoding. Instead of matching the original model output distribution token-by-token, we identify the generated tokens that affect the downstream quality of the response, relaxing the distribution match guarantee so that the \"unimportant\" tokens can be generated faster. Our approach relies on a semi‑greedy search algorithm to test which of the mismatches between target and draft models should be corrected to preserve quality and which ones may be skipped. We then train a lightweight classifier based on existing LLM embeddings to predict, at inference time, which mismatching tokens can be safely accepted without compromising the final answer quality. We evaluate AutoJudge with multiple draft/target model pairs on mathematical reasoning and programming benchmarks, achieving significant speedups at the cost of a minor accuracy reduction. Notably, on GSM8K with the Llama 3.1 70B target model, our approach achieves up to \r\n≈\r\n2\r\n×\r\n speedup \\textit{over speculative decoding} at the cost of a \r\n≤\r\n1\r\n%\r\n drop in accuracy. When applied to the LiveCodeBench benchmark, AutoJudge automatically detects programming-specific important tokens, accepting \r\n≥\r\n25\r\n tokens per speculation cycle at a\r\n \r\n2\r\n%\r\n drop in Pass@1. Our approach requires no human annotation and is easy to integrate with modern LLM inference frameworks."}],"_id":"22830","scopus_import":"1","date_published":"2025-12-02T00:00:00Z","quality_controlled":"1","date_updated":"2026-09-10T11:08:33Z","oa":1,"page":"104904-104941","status":"public","volume":38,"title":"AutoJudge: Judge decoding without manual annotation","arxiv":1,"month":"12","external_id":{"arxiv":["2504.20039"]},"intvolume":" 38","oa_version":"Published Version","day":"02","year":"2025","department":[{"_id":"DaAl"}],"type":"conference","publisher":"Neural Information Processing Systems Foundation","doi":"10.52202/085713-3163","alternative_title":["Advances in Neural Information Processing Systems"],"publication_status":"published","language":[{"iso":"eng"}],"user_id":"2DF688A6-F248-11E8-B48F-1D18A9856A87","citation":{"mla":"Garipov, Roman, et al. “AutoJudge: Judge Decoding without Manual Annotation.” 39th Annual Conference on Neural Information Processing Systems, vol. 38, Neural Information Processing Systems Foundation, 2025, pp. 104904–41, doi:10.52202/085713-3163.","ama":"Garipov R, Velikonivtsev F, Ermakov I, Svirschevski R, Egiazarian V, Ryabinin M. AutoJudge: Judge decoding without manual annotation. In: 39th Annual Conference on Neural Information Processing Systems. Vol 38. Neural Information Processing Systems Foundation; 2025:104904-104941. doi:10.52202/085713-3163","ieee":"R. Garipov, F. Velikonivtsev, I. Ermakov, R. Svirschevski, V. Egiazarian, and M. Ryabinin, “AutoJudge: Judge decoding without manual annotation,” in 39th Annual Conference on Neural Information Processing Systems, San Diego, CA, United States, 2025, vol. 38, pp. 104904–104941.","apa":"Garipov, R., Velikonivtsev, F., Ermakov, I., Svirschevski, R., Egiazarian, V., & Ryabinin, M. (2025). AutoJudge: Judge decoding without manual annotation. In 39th Annual Conference on Neural Information Processing Systems (Vol. 38, pp. 104904–104941). San Diego, CA, United States: Neural Information Processing Systems Foundation. https://doi.org/10.52202/085713-3163","short":"R. Garipov, F. Velikonivtsev, I. Ermakov, R. Svirschevski, V. Egiazarian, M. Ryabinin, in:, 39th Annual Conference on Neural Information Processing Systems, Neural Information Processing Systems Foundation, 2025, pp. 104904–104941.","chicago":"Garipov, Roman, Fedor Velikonivtsev, Ivan Ermakov, Ruslan Svirschevski, Vage Egiazarian, and Max Ryabinin. “AutoJudge: Judge Decoding without Manual Annotation.” In 39th Annual Conference on Neural Information Processing Systems, 38:104904–41. Neural Information Processing Systems Foundation, 2025. https://doi.org/10.52202/085713-3163.","ista":"Garipov R, Velikonivtsev F, Ermakov I, Svirschevski R, Egiazarian V, Ryabinin M. 2025. AutoJudge: Judge decoding without manual annotation. 39th Annual Conference on Neural Information Processing Systems. NeurIPS: Neural Information Processing Systems, Advances in Neural Information Processing Systems, vol. 38, 104904–104941."},"conference":{"name":"NeurIPS: Neural Information Processing Systems","end_date":"2025-12-07","start_date":"2025-12-02","location":"San Diego, CA, United States"},"publication":"39th Annual Conference on Neural Information Processing Systems","article_processing_charge":"No"}