Standard modularity is unsuitable for functional regionalization of spatial interaction data
Bibliographic Data
| ID | 12401040 |
|---|---|
| Authors | Lucas Martínez‐Bernabéu (0000-0002-9532-8038, University of Alicante, corresponding author), José María Casado (0000-0001-5310-7678, University of Alicante) |
| Year | 2021 |
| Volume | 100 |
| Issue | 5 |
| Pages | 1323-1331 |
| Publication date | 2021-05-14 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Papers of the Regional Science Association (JOURNAL) |
| Journal identifiers | ISSN: 1056-8190 • E-ISSN: 1435-5957 |
| Publisher | Elsevier BV (PUBLISHER) |
| DOI | 10.1111/pirs.12617 |
| OpenAlex | W3162861216 |
| Language | EN |
| Citations received | 2 |
| References cited | 14 |
Functional regions capturing local socioeconomic dynamics are a framework for territorial statistics and policy-making. The usefulness of such statistics and policies depends on the adequateness of the regions used in the analysis. Several authors have adopted the modularity quality function to drive their regionalization methods. This function, originally devised for non-spatial data, has limitations particularly relevant in this context. The paper discusses these limitations and illustrate them using real data for the United States. The results warn against using standard modularity to assess or optimize the quality of a functional delineation, and advocate for available alternatives that consider the spatial nature of functional regionalization. Las regiones funcionales que captan la dinámica socioeconómica local son un marco para las estadísticas y la elaboración de políticas territoriales. La utilidad de estas estadísticas y políticas depende de lo adecuado de las regiones utilizadas en el análisis. Varios autores han adoptado la función de calidad de la modularidad para potenciar sus métodos de regionalización. Esta función, concebida originalmente para datos no espaciales, tiene limitaciones especialmente relevantes en este contexto. Este artículo analiza estas limitaciones y las ilustra mediante el uso de datos reales de Estados Unidos. Los resultados desaconsejan el uso de la modularidad estándar para evaluar u optimizar la calidad de una delimitación funcional, y abogan por las alternativas disponibles que tienen en cuenta la naturaleza espacial de la regionalización funcional. 地域社会経済の動態を捉える機能地域(functional regions)は、地域統計と政策決定のフレームワークである。これらの統計や政策の有用性は、分析の対象となった地域の適切性に依存する。研究者には、各自の地域区分法を補強するためにモジュラリティ品質関数を採用している者もいる。この関数は、もともと非空間データ用に考案されたもので、このコンテキストでは特に重要な限界がある。本稿では、これらの限界を考察し、米国の実際のデータを用いて検証する。結果から、機能区分の質を評価または最適化するために標準的なモジュラリティを使用しないよう警告が与えられ、地域区分の空間的性質を考慮する利用可能な選択肢を用いることが提唱される。 The definition of appropriate functional regions (FR) is crucial for carrying out meaningful analyses of socio-economic phenomena, and for the design, implementation and monitoring of regional public policies. In the last decades a variety of regionalization methods have been developed and applied both in the academic and administrative spheres in order to improve the quality of these areas (Casado-Díaz & Coombes, 2011; OECD, 2002, 2020; Van der Laan & Schalke, 2001). In recent years, many authors have adopted modularity quality function Q to drive the regionalization processes to delimit FRs based on commuting flows or mobile phone data. These are the cases of the labour market areas (LMAs) defined by De Montis et al. (2013) in Sardinia, Farmer and Fotheringham (2011) in Ireland, Kropp and Schwengler (2016) in Germany, and Shen and Batty (2019) in London's Metropolitan Area. Nelson and Rae (2016) compare the results of a modularity optimization algorithm and a visual heuristic to delimit economic areas in the US. In some works modularity is the main criterion used to assess the quality of regionalizations (e.g. Amini et al., 2014, measure the adequacy of several methods applied in Portugal and Ivory Coast; Wicht et al., 2020, compare regionalizations in Germany). Modularity was originally designed in the field of network analysis and community detection (Newman & Girvan, 2004). In short, standard modularity compares the share of edges falling within each community (group of nodes) with the expected share if the edges from each node were distributed among all nodes proportionally to the nodes' sizes. When the weight of the edges between two communities exceeds the expected one, merging such communities increases total modularity score. The use of modularity in its original context has a well-known limitation, the resolution limit (Fortunato & Barthelemy, 2007; Lancichinetti & Fortunato, 2011): for large (small) enough networks, the expected number of edges between two communities is smaller (larger) than the weight of most edges in the network. When this happens, even a weak (strong) interaction between two objectively unrelated (related) communities is interpreted by modularity as a sign of a strong (weak) interaction, merging them (keeping separated). “Modularity optimization has an intrinsic bias towards partitions having a characteristic number of modules which might not be compatible with the modular organization of the system” (Lambiotte, 2010, p. 546). In the field of spatial regionalization, this limitation is seldom mentioned. Sometimes it is negated (e.g., “the modularity measure Q is unbiased with regard to the number of defined regions”, Kropp & Schwengler, 2016, p. 430; this statement is false, as we show in this paper). Usually, authors perform no comparison to results from other established methods, and greater modularity scores are taken as signs of an improved regionalization without much discussion. Only a few state the problem (e.g., Nelson & Rae, 2016) or attempt to address it explicitly through modifications in the procedure (see Section 3). We argue the shortcomings of modularity are critical in spatial functional regionalization and advocate for using other quality indicators. In Section 2 we explain standard modularity and illustrate the relevance of its drawbacks using a formulaic proof and real-world case studies. In section 3 we highlight works that proposed adaptations of modularity to the spatial context. Section 4 concludes. In other words, when the aggregated flow between a and b divided by the product of their sizes is greater than an arbitrary constant, which depends on the size of the territory under analysis and that is unrelated to the functional characteristics of the regions. To illustrate the strong dependence of modularity scores on the ratio FR's size to territory's size, Table 1 shows the merge thresholds for two scenarios with different relative size of the regions under consideration, Tr, with respect to the whole territory, T, each scenario sub-divided into three cases of region size relative to each other. In scenario A, the flows between regions expected by modularity are so high that strongly dependent regions are kept separated, producing insufficiently autonomous regions (e.g., a micro or small region sharing 49% or 44% of its workers would not be absorbed by the larger region). In scenario B, negligible interactions (< 1%) trigger merging areas, creating too large and insufficiently cohesive regions. Under the assumption that the ideal number of LMAs depends on the functional relationships between the BTUs, the regionalizations identified for the whole USA in each case should be similar. However, the results do not meet our expectations (Table 2). Conterminous US is divided into 78 LMAs when regionalized alone (c.1) or together with Alaska and Hawaii (a.1). When regionalized separated into the four Census Bureau-designated regions (b.1–4), Conterminous US totals 171 LMAs. Within the Double US experiment, it gets 56 LMAs. The effects in Alaska and Hawaii are stronger due to bigger differences in the region-territory size ratio. Alaska (with 27 counties/census areas) is divided in 22 LMAs when regionalized alone (c.2), two LMAs within the West region (b.4), a single Alaskan LMA within Full US (a.1). Hawaii, with five counties: five LMAs if regionalized alone, three LMAs if regionalized with the West region, and a single LMA if regionalized within Full or Double US. In Wyoming (the least populous state in Conterminous US), with 23 counties, when regionalized within the West region, four counties form a purely Wyoming LMA with 135,000 inhabitants and 94.46% autonomy (without further analysis of its cohesion, it seems a properly sized and very autonomous LMA) and the remaining 19 counties are allocated to three other LMAs from two different states. This LMA disappears in a.1 and c.1 results, its counties allocated to two LMAs (one covering most of Colorado and the other covering western South Dakota and Nebraska). Those LMAs in Western US in a.1/c.1 are too big to be cohesive and cannot be considered LMAs but actually composed of independent LMAs fused together. The fact that Wyoming LMA is divided into two in a.1/c.1 (instead of all its counties being absorbed by a single external LMA) suggests that there might be a division of it into two, more cohesive LMAs in b.4, or that its dismembering in a.1/c.1 is not ideal because those four counties should be kept together. If one of these US regionalizations represents the proper LMAs, the others are necessarily far from ideal. In any case, modularity does not indicate which is the right regionalization, it is the users who must find it out by examining the inter-BTU flows within and between LMAs based on other indicators (e.g., Martínez-Bernabéu et al., 2020) and their own knowledge. Greater values of λ produce larger FRs. However, as Lambiotte points out, introducing a resolution parameter lacks theoretical ground, and still requires expert knowledge and the repetition of the method with different values of λ, in a trial-and-error procedure with “the same limitations as modularity when the resolution parameter is fixed [...] additional tests are therefore needed to uncover significant scales of description” (Lambiotte, 2010, p. 552). This can be used to uncover a hierarchy of nested regionalizations, to let the user choose the right one, but still it “is not able to uncover coarser partitions than those obtained by modularity optimization. Moreover, it may produce hierarchies even when the system is single-scale or, worse, completely random” (Lambiotte, 2010, p. 549). There is no guarantee that all relevant FRs are identified, because modularity “suffers from two opposite coexisting problems: the tendency to merge small subgraphs, which dominates when the resolution is low; the tendency to split large subgraphs, which dominates when the resolution is high [...] the simultaneous elimination of both biases is not possible and multiresolution modularity is not capable to recover the planted community structure” (Lancichinetti & Fortunato, 2011, p. 1). A synthetic example can illustrate this: four small counties with 55% local autonomy (too low to be considered LMAs) send 9% of their workforce to each of the other three; another two large counties have 90% local autonomy (arguably good LMAs by their own) and send the remaining 10% of their workforce to the other large county. A merging threshold lower than 9% would allow the merge of the four small counties into a single LMA with 82% local autonomy and 27% of inter-BTU flows (arguably a good LMA in terms of autonomy and cohesion). However, to keep separated the two large counties the threshold should be higher than 10%. No single threshold will allow to identify all relevant FRs. Moreover, even when a single threshold suffices, the results obtained might make no sense in a spatially constrained context, for BTUs that do not show a clear pattern of interaction with any region (“Because these weakly-linked nodes could be assigned into almost any different community with little result on the achieved modularity score, these nodes often caused trouble with geographic coherence, as the algorithm assigned them to far-flung communities which made little sense from an interpretive standpoint” Nelson & Rae, 2016, p. 12). The only way to guarantee high quality while using standard modularity is to apply different thresholds to different parts of the territory, but modularity does not provide any tool to deduce the right thresholds or to apply different ones on different parts of the territory. A method with sophisticated graphical interface could allow to choose different thresholds for different BTUs or regions (or branches of the dendrogram of regionalizations) and plot the results to let the user adjust the values. When the desired scale of division is reached in all areas, the interface could highlight regions or BTUs with bad values in relevant indicators to consider reallocations of BTUs between regions. Anyway, it is still up to the user to decide the correct regionalization based on their expert knowledge, the method could not provide a good regionalization without such expert knowledge. This is undesirable in situations where previous experience is not available. Some authors have proposed improvements to the modularity equation to reduce or eliminate the dependency on the territory size. Farmer and Fotheringham (2011) use a Gaussian-type inverse distance weighting scheme to adjust the values of commuting flows depending on distance between origin and destination, with a parameter to control the bandwidth of the Gaussian operator. Expert et al. (2011), Fukumoto et al. (2014), Gao et al. (2013) and Liuet al. (2014) substitute modularity's null model for a classic gravity model to estimate expected flows within each FR. A distance-decay parameter is obtained by fitting the gravity model to the observed commuting flows. Gao et al. (2013) also considers a fraction (instead of difference) format modularity to better represent the differences of scale between small/sparsely and large/densely populated regions. Fukumoto et al. (2014) noted that the method is not robust to the choice of the time decaying parameter, which could be an issue to get comparable regions in different territories. Although the results of these alternatives also require fixing parameters or a final step of manual adjustments, the resulting regionalizations are arguably sound and robust, the foundations of their equations make sense in a spatially constrained context, the parameters are readily understood (e.g., minimum autonomy of an LMA), and they do not require the user to examine dendrograms of regionalizations or to perform extensive changes. Any of them can achieve results better than standard modularity, and require less knowledge and man-hours than multi-resolution modularity. Given the growing relevance of regional analysis for public policy, and the potentially serious implications of using inappropriate FRs, practitioners should acknowledge the limitations of the available approaches and use the regionalization method that better suits their needs. The utility of a quality function is to determine the (near) optimal solution. Modularity as a global fitness function is not useful to determine the desired scale of division, or to assess the adequateness of the allocation of individual BTUs. That is left to the practitioner, who has to navigate through alternative results and manually tailor them based on expert knowledge. Hence, modularity is of much less usefulness than it is often given credit for, and methods based on its maximization are therefore inappropriate for policy-makers. Many works carried out in this area in recent years have ignored this limitations. Practitioners should be aware of the available alternatives to standard modularity that provide sound and comparable results for territories of any size. Our research agenda includes performing a comparative analysis of these and other non-modularity quality functions in different optimization methods for different datasets, to identify the best methods or points of improvement. But we already know that any of those is a better choice than standard modularity. This work was funded by the Spanish Ministry of Science, Innovation and Universities/Agencia Estatal de Investigación (AEI) and the European Regional Development Fund, grant number CSO2017-86474-R. The funding source had no involvement in study design; in the collection, analysis and interpretation of data; in the writing of the report; or in the decision to submit the article for publication
Context (archaeology · Data mining · Data science · Econometrics · Function (biology · Geography · Modularity (biology · Quality (philosophy · Spatial analysis · Statistics · Computer Science · Land Use and Ecosystem Services · Mathematics · Regional Economics and Spatial Analysis · Spatial and Panel Data Analysis
Discovering Spatial Interaction Communities from Mobile Phone D ata
Resolution limit in community detection
Finding and evaluating community structure in networks
Are functional regions more homogeneous than administrative regions? A test using hierarchical linear models
Three-Step Method for Delineating Functional Labour Market Regions
Delineating the perceived functional regions of London from commuting flows
Functional Regionalisation of Spatial Interaction Data. An Evaluation of Some Suggested Strategies
Network-Based Functional Regions
| Unique citing works | 2 |
|---|---|
| Citations per year | 0,67 |
| Citation span | 2023 - 2026 (4) |
| Citation velocity | current |
| Highly cited | No |
| Citation types | Neutral: 2 |