Hello, thank you for your great work.
I am currently working on reproducing the results for LongViLa-R1 7B on the VideoMME benchmark.
I noticed that the paper reports a score of 65.1 on VideoMME (without subtitles). Could you please clarify if this performance was achieved using "Think Mode"?
I would appreciate it if you could share the specific inference settings used to achieve this score.
Thank you in advance!
Hello, thank you for your great work.
I am currently working on reproducing the results for LongViLa-R1 7B on the VideoMME benchmark.
I noticed that the paper reports a score of 65.1 on VideoMME (without subtitles). Could you please clarify if this performance was achieved using "Think Mode"?
I would appreciate it if you could share the specific inference settings used to achieve this score.
Thank you in advance!