MapTab: A Diagnostic Benchmark for Long-HorizonMulti-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs

Ziqiao Shang1,2†, Ling-Yue Ge1,2†, Zian Xu1,2, Zi-Jian Cheng1,2,
Shi-Yu Tian1,2, Zhenyu Huang1,2, Wenbo Fu1,2, Weiming Wu1,2,
Yang Chen1,2, Xiangwen Zhang3, Yulan Hu3, Bin Liu4, Lan-Zhe Guo1,2*
Submitted to AAAI 2027
1National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China 2School of Intelligence Science and Technology, Nanjing University, Suzhou, China 3AMAP, Alibaba Group 4School of Computing and Artificial Intelligence, Southwest Jiaotong University, Chengdu, China
Equal contribution*Corresponding author
Correspondence: guolz@lamda.nju.edu.cn
Nanjing University
LAMDA
AMAP
Southwest Jiaotong University
MapTab overview
MapTab is a comprehensive benchmark designed to evaluate the map understanding and spatial reasoning capabilities of Vision-Language Models (VLMs). The benchmark focuses on two core tasks: route planning and map-based question answering, using both metro maps and travel maps.

Abstract

Systematically evaluating Multimodal Large Language Models (MLLMs) is essential for advancing Artificial General Intelligence (AGI). Yet existing benchmarks remain inadequate for rigorously measuring their reasoning capabilities under multi-criteria constraints. To address this gap, we introduce MapTab, a multimodal benchmark designed to assess holistic multi-criteria reasoning in MLLMs through route-planning tasks. MapTab requires models to perceive and ground visual information from map images while integrating route attributes, such as Time and Price, from structured tables. It covers two scenarios: Metromap, spanning metro networks in 160 cities across 52 countries, and Travelmap, featuring 168 representative tourist attractions from 19 countries. Overall, MapTab includes 328 images, 196,800 route-planning queries, and 3,936 QA queries, incorporating four key criteria: Time, Price, Comfort, and Reliability. Extensive evaluations of 21 representative MLLMs show that current models still struggle with multicriteria multimodal reasoning. Notably, when visual perception is unreliable, multimodal reasoning can even underperform unimodal approaches. MapTab therefore offers a challenging and realistic testbed for systematically evaluating and advancing MLLMs across core perception, integration, numerical comparison, and route planning capabilities.

Route planning leaderboard

The following leaderboard presents the evaluation results of various Multimodal Large Language Models (MLLMs) across different data modalities in the MapTab path planning task. Performance is measured using three key metrics: Exact Match Accuracy (EMA), Partial Match Accuracy (PMA), and Difficulty-aware Score (DS). Models were tested using varying combinations of map, edge, and vertex data, as detailed below:
—Map-only: Only map data used
—Edge-only: Only edge data used
—Map+Edge: Map and edge data combined
—Map+Edge+Vertex: Map, edge, and vertex data combined
—Map+Mix: Map and merged tabular data (Mix_tab)
For clarity, comparisons involving Edge_tab + Vertex_tab were omitted, as they yielded similar results to the Map-only and Edge_tab-only groups without adding new insights. The best performing results in both open-source and closed-source model groups are highlighted in bold.

ModelType Map-only Edge-only Map+Edge Map+Edge+Vertex Map+Mix
EMAPMADS EMAPMADS EMAPMADS EMAPMADS EMAPMADS
Scenario: MetroMap
Open-source Models
Qwen3-VL-8B-InstructNo-Thinking2.7517.5810325.6946.44115321.2541.3092119.3139.318554.6921.87182
Qwen3-VL-8B-ThinkingThinking5.1220.9918831.6949.76142738.0057.06177123.7541.6910806.3822.93270
Qwen3-VL-2B-InstructNo-Thinking0.9415.14359.8827.614376.6323.852827.0026.913252.0017.8278
Qwen2.5-VL-7B-InstructNo-Thinking0.9415.023214.0031.2062111.6928.325087.9420.773573.3818.09131
Phi-3.5-Vision-Instruct-4BNo-Thinking0.0610.40210.8727.924766.6322.142722.7512.271170.8112.9426
Phi-4-Multimodal-Instruct-6BNo-Thinking0.009.7502.1312.52842.1311.78911.759.51700.449.0214
InternVL3-8B-InstructNo-Thinking0.1313.98410.5029.5746012.8131.835559.0024.734131.7517.0076
Qwen3-VL-30B-A3B-InstructNo-Thinking3.3119.2612923.6944.33106222.5643.58101719.0040.038426.7526.22288
Qwen3-VL-32B-InstructNo-Thinking6.3122.2325031.8754.45143032.1254.54146328.5050.0613036.5624.43262
Qwen3-VL-32B-ThinkingThinking13.3129.4355831.8154.94145144.1262.77214826.5651.4812119.1928.89381
Qwen3.5-9BNo-Thinking5.6922.4423125.5648.95117426.5049.33118920.2544.159038.7528.15365
Qwen3.6-35B-A3BThinking11.5629.1647664.8877.06351053.8771.43281143.6964.16227514.7534.76643
Closed-source Models
GPT4-oNo-Thinking6.6325.6125742.3864.07209840.6962.40196935.6355.51170211.3131.11469
GPT4.1No-Thinking7.9425.5230648.5667.07244646.8165.18234441.8162.88205914.0635.98608
GPT-5.5-InstantNo-Thinking43.6364.27232088.7596.24544185.3193.89524488.8892.15546769.5079.764021
Doubao-Seed-1.6-w/o-ThinkingNo-Thinking8.1324.6031546.9466.98235148.0666.95243440.5662.11204113.8135.61579
Doubao-Seed-1.6-ThinkingThinking12.0630.4951274.3886.23428474.0085.68425676.0683.41435622.0342.481029
Qwen-VL-Plus-w/oThinkingNo-Thinking4.8121.8318636.8858.69172538.2558.59180431.6252.9214796.9427.69288
Qwen-VL-Plus-ThinkingThinking10.7529.1143761.5076.62327662.1976.42333145.7564.46229016.3837.44714
Gemini-3-Flash-PreviewNo-Thinking37.0657.15188174.7584.99445773.0683.37431669.1976.14406053.8765.842976
Gemini-3.5-flashNo-Thinking44.1965.40243583.8793.53512583.7593.04509982.3785.34499360.5071.673433
Scenario: TravelMap
Open-source Models
Qwen3-VL-8B-InstructNo-Thinking19.2942.50104044.0561.66259743.3361.39254034.5255.56200215.6540.97804
Qwen3-VL-8B-ThinkingThinking22.6245.94120374.1782.41451182.6888.54519933.1555.60185712.7438.10653
Qwen3-VL-2B-InstructNo-Thinking8.4534.3045011.2532.3564419.1745.68106612.1440.476893.1530.69159
Qwen2.5-VL-7B-InstructNo-Thinking7.6830.4839021.0738.15116024.8242.02133715.6037.478184.7028.84228
Phi-3.5-Vision-Instruct-4BNo-Thinking0.1220.00612.2034.816819.8231.875434.4623.212341.3122.6867
Phi-4-Multimodal-Instruct-6BNo-Thinking0.4219.26207.2017.634105.3015.932831.739.36961.4318.9667
InternVL3-8B-InstructNo-Thinking6.6129.2130829.5849.69163629.4050.16164713.5736.787992.5024.28132
Qwen3-VL-30B-A3B-InstructNo-Thinking17.8644.1597650.9565.36297853.7567.71318638.4558.0223189.7037.93517
Qwen3-VL-32B-InstructNo-Thinking36.9057.44209164.5276.16394468.3978.99426952.5669.18316921.6747.341161
Qwen3-VL-32B-ThinkingThinking39.1758.84226269.7679.60430791.7994.55594942.3262.99249219.9446.731082
Qwen3.5-9BNo-Thinking25.9549.39142440.8354.74248750.0666.84293933.9357.12192316.4942.32892
Qwen3.6-35B-A3BThinking42.7461.74246195.4296.93628992.2694.68601762.0274.68382527.0849.281532
Closed-source Models
GPT4-oNo-Thinking16.8540.9893065.0675.84465162.7474.11446746.0763.07306912.0838.07675
GPT4.1No-Thinking20.3043.24107774.8282.98465070.8979.84436454.7069.59331015.0640.67793
GPT-5.5-InstantNo-Thinking68.1578.51424298.5199.06659398.0498.55654683.9987.04551057.9869.873535
Doubao-Seed-1.6-w/o-ThinkingNo-Thinking33.0454.15188073.5182.16455276.8584.04481956.2571.46339025.4849.541406
Doubao-Seed-1.6-ThinkingThinking38.4558.46229598.3998.87656797.8698.47652783.1589.08542325.3048.901437
Qwen-VL-Plus-w/o-ThinkingNo-Thinking30.6052.64169164.2376.45392169.6479.78428953.9970.07324322.9247.651262
Qwen-VL-Plus-ThinkingThinking38.2758.94219464.3576.53392994.2396.04615956.1970.84340823.2147.181289
Gemini-3-Flash-PreviewNo-Thinking60.0073.20369398.2798.38656794.4094.87625378.5182.40519443.5160.112687
Gemini-3.5-flashNo-Thinking68.2179.12418699.5299.69667498.5798.95659992.5094.28617758.9371.303648

QA leaderboard

This leaderboard summarizes the performance of various Multimodal Large Language Models (MLLMs) on QA tasks across the MetroMap and TravelMap scenarios. Input modalities are represented as:
—M for Map
—E for Edge_tab
—V for Vertex_tab
In the MetroMap scenario, the Mix_tab paired with the Map input excludes the Line column to minimize unnecessary table details, ensuring the evaluation focuses on map-table coordination.
Tasks are categorized into three distinct types:
—Global Perception-based Reasoning Tasks (GP)
—Local Perception-based Reasoning Tasks (LP)
—Spatial Relationship Judgment Tasks (SR)
Bold values in the table indicate the best performance for open-source and closed-source models, respectively.

ModelType Map (M) Edge (E) Vertex (V) Map+Mix_tab
GPLPSR GPLPSR GPLPSR GPLPSR
Scenario: MetroMap
Open-source Models
Qwen3-VL-8B-InstructNo-Thinking55.0017.5073.1222.50100.07.5057.5051.8886.880.6322.5038.75
Qwen3-VL-8B-ThinkingThinking53.1228.1251.2556.8799.3856.8779.3777.5098.127.509.3835.63
Qwen3-VL-2B-InstructNo-Thinking8.135.0063.123.7587.503.1226.2511.2564.380.008.1326.87
Qwen2.5-VL-7B-InstructNo-Thinking48.7515.6266.2515.62100.010.0044.3760.6287.501.2525.6233.12
Phi-3.5-Vision-Instruct-4BNo-Thinking58.1318.7578.1222.50100.021.8853.7566.8798.120.6321.2540.00
Phi-4-Multimodal-Instruct-6BNo-Thinking60.6240.6280.0041.2593.1342.5058.7592.5099.385.6335.6349.38
InternVL3-8B-InstructNo-Thinking35.0020.6265.007.5083.133.1233.7510.0070.630.0012.5028.12
Qwen3-VL-30B-A3B-InstructNo-Thinking20.6216.2560.620.6368.751.250.001.2568.750.0011.2518.12
Qwen3-VL-32B-InstructNo-Thinking5.000.0038.1210.6278.123.1216.254.3773.120.007.5023.75
Qwen3-VL-32B-ThinkingThinking26.8719.3850.005.0099.385.0021.2525.0085.620.006.8860.62
Qwen3.5-9BNo-Thinking59.3817.5076.255.6399.388.7557.5061.2576.251.2520.0075.62
Qwen3.6-35B-A3BThinking60.6242.5082.5071.8898.7583.1385.6298.12100.015.0045.6295.63
Closed-source Models
GPT4-oNo-Thinking62.5013.7578.7531.87100.028.1255.6378.12100.03.7529.3845.00
GPT4.1No-Thinking61.8826.8776.2550.6299.3838.1264.3883.75100.03.1225.6249.38
GPT-5.5-InstantNo-Thinking58.7583.1391.2599.38100.096.25100.0100.0100.068.7581.8796.88
Doubao-Seed-1-6-251015-w/o_ThinkingNo-Thinking55.6320.6276.2541.88100.059.3858.1386.2599.383.7549.3850.62
Doubao-Seed-1-6-251015-ThinkingThinking54.3740.6277.5072.50100.069.3796.2598.75100.027.5050.0053.12
Qwen-VL-Plus-w/o_ThinkingNo-Thinking60.0021.8878.7540.62100.040.0064.3877.5097.501.2525.6240.62
Qwen-VL-Plus-ThinkingThinking57.5045.0081.8768.75100.071.2590.6295.63100.013.7546.2555.63
Gemini-3-Flash-PreviewNo-Thinking59.3882.5093.1391.2598.1275.6288.7594.37100.048.1380.0094.37
Gemini-3.5-flashNo-Thinking63.7586.2588.75100.099.3898.7598.1299.3899.3888.7582.5096.88
Scenario: TravelMap
Open-source Models
Qwen3-VL-8B-InstructNo-Thinking7.1460.1252.9817.8699.4045.2438.6950.6061.3175.6070.2414.29
Qwen3-VL-8B-ThinkingThinking39.2970.8352.3887.50100.039.88100.0100.0100.063.1069.0513.69
Qwen3-VL-2B-InstructNo-Thinking12.5058.939.526.0094.6464.881.1946.4338.1064.2967.864.17
Qwen2.5-VL-7B-InstructNo-Thinking4.7665.4854.1738.1099.4047.6217.2687.5079.7633.9368.4516.07
Phi-3.5-Vision-Instruct-4BNo-Thinking13.1048.2159.5239.8899.4041.0750.0075.6077.3872.0274.4010.71
Phi-4-Multimodal-Instruct-6BNo-Thinking44.0570.2458.9378.5798.8133.9396.4398.21100.073.2169.0517.86
InternVL3-8B-InstructNo-Thinking12.5059.5235.718.9397.6249.4010.7150.0051.7970.8367.864.76
Qwen3-VL-30B-A3B-InstructNo-Thinking5.9550.606.557.7486.3152.388.9344.6445.2454.7635.122.98
Qwen3-VL-32B-InstructNo-Thinking0.0042.2614.8811.3163.6938.3118.4545.8347.6227.9863.695.95
Qwen3-VL-32B-ThinkingThinking8.3360.1223.211.1997.0244.6410.1238.6956.5569.0566.675.95
Qwen3.5-9BNo-Thinking25.0066.0754.768.3399.4069.6420.2461.3175.6082.7471.4321.43
Qwen3.6-35B-A3BThinking23.2175.0064.2997.62100.077.38100.0100.0100.079.7676.1925.00
Closed-source Models
GPT4-oNo-Thinking11.3163.6949.4047.02100.036.3153.5799.4067.2671.4373.2111.31
GPT4.1No-Thinking3.5769.6455.9547.02100.042.8666.07100.069.6473.2177.3817.86
GPT-5.5-InstantNo-Thinking44.0591.6782.14100.099.4098.81100.0100.0100.080.3697.6232.74
Doubao-Seed-1-6-251015-w/o_ThinkingNo-Thinking22.6258.9350.0063.1098.8151.7954.76100.098.8166.0776.7925.60
Doubao-Seed-1-6-251015-ThinkingThinking24.4071.4355.9595.8384.5271.4397.62100.0100.078.5772.0222.02
Qwen-VL-Plus-w/o_ThinkingNo-Thinking19.6472.0252.9848.81100.035.7157.7498.8182.7454.1770.2419.64
Qwen-VL-Plus-ThinkingThinking24.4067.8662.5098.81100.034.52100.0100.0100.069.0572.6224.40
Gemini-3-Flash-PreviewNo-Thinking45.8385.1277.9897.6299.4086.31100.099.40100.085.7181.5526.19
Gemini-3.5-flashNo-Thinking60.1289.8876.19100.099.4093.4599.40100.0100.082.7495.8329.76

BibTeX

@article{shang2026maptab,
  title=MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs},
  author={Shang, Ziqiao and Ge, Ling-Yue and and Xu, Zian and Cheng, Zi-Jian and Tian, Shi-Yu and Huang, Zhenyu and Fu, Wenbo and Wu, Weiming and Chen, Yang and Zhang, Xiangwen and Hu, Yulan and Bin, Liu and Guo, Lan-Zhe},
  journal={arXiv preprint arXiv:2602.18600},
  year={2026}
}