Hi team,
First of all, thank you for releasing this benchmark and the sample inference code—it’s been incredibly helpful.
I’m currently trying to benchmark some methods on REPOCOD but am encountering difficulties reproducing the results. Specifically, I’m using VLLM for generation, and the outputs differ from those produced via direct inference using HuggingFace Transformers.
Here’s a snippet of the inference code I’m using with VLLM:
model_path = 'deepseek-coder-6.7b-base'
tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
custom_stop_token = tokenizer.encode("```", add_special_tokens=False)[0]
sampling_params = SamplingParams(temperature=1.0, top_k=50, top_p=1.0, max_tokens=4096, stop_token_ids=[custom_stop_token])
llm = LLM(model=model_path, dtype=torch.float16, max_model_len=16_384, enforce_eager=True)
current_file_prompt = current_file_template.format(sample['target_module_path'], prefix, suffix, sample['prompt'])
input_text = f"{SYSTEM_PROMPT}\n{current_file_prompt}"
output = llm.generate([input_text], sampling_params)
Could you please:
- Share the inference code you used with VLLM?
- Suggest any modifications or configurations I might have overlooked?
Hi team,
First of all, thank you for releasing this benchmark and the sample inference code—it’s been incredibly helpful.
I’m currently trying to benchmark some methods on REPOCOD but am encountering difficulties reproducing the results. Specifically, I’m using VLLM for generation, and the outputs differ from those produced via direct inference using HuggingFace Transformers.
Here’s a snippet of the inference code I’m using with VLLM:
Could you please: