Hi, thank you for your fabulous project.
While I'm not using slurm and S3 when running evaluation, I have several questions about lmms-eval_videochat/lmms_eval/models/qwen2_5_vl_lxh.py.
1. The total pixels is hardcoded at line 219:
message.append({"role": "user", "content": [{"type": "video", "video": video_path, "media_dict": media_dict, "total_pixels": 3584 * 28 * 28}, {"type": "text", "text": context}]})
Which will result 2,812,928 pixels, while the model is initiated to be 1,605,632 pixels.
If you can clarify the max_pixels used in you experiment, that will be appreciated.
2.
It seems this model class doesn't limit the number of video frames as Qwen2.5VL did, will this contribute to OOM Error?
3.
It seems we are assuming each batch only contains one video?
Thank you!
Hi, thank you for your fabulous project.
While I'm not using slurm and S3 when running evaluation, I have several questions about lmms-eval_videochat/lmms_eval/models/qwen2_5_vl_lxh.py.
1. The total pixels is hardcoded at line 219:
Which will result 2,812,928 pixels, while the model is initiated to be 1,605,632 pixels.
If you can clarify the max_pixels used in you experiment, that will be appreciated.
2.
It seems this model class doesn't limit the number of video frames as Qwen2.5VL did, will this contribute to OOM Error?
3.
It seems we are assuming each batch only contains one video?
Thank you!