The Clustering and Co-Organization analysis tool aims to give insights into the spatial organization of diverse molecular distributions in a quantitative and automatic manner.
Overview
The Clustering and Co-Organization analysis tool is based on the Statistical Object Distance Analysis (SODA) algorithm (Lagache et al. 2018), a spatial statistics method used to analyze whether one or more molecular populations are organized together more frequently than would be expected in a random spatial distribution.
Unlike traditional colocalization methods that rely on pixel overlap or fluorescence intensity correlation, SODA analyses the spatial relationships between detected molecular objects (clusters of localizations) using distance-based statistics. Specifically, SODA evaluates whether two populations of localization clusters are spatially associated beyond a random distribution. Importantly, SODA does not directly prove biochemical binding, physical interaction, or exact molecular overlap. Instead, it identifies statistically significant or enriched spatial organization events.
In addition, traditional colocalization metrics, such as Pearson Correlation Coefficient or Manders Overlap Coefficient, strongly depend on fluorescent intensity and molecular density, which work poorly in dense biological structures, as molecules may appear highly overlapping even when randomly distributed.
SODA addresses this limitation by:
- Analyzing object positions directly
- Accounting for molecular density
- Correcting for overlapping neighborhoods
This makes SODA particularly well suited for super-resolution and localization microscopy datasets.
How SODA works
SODA operates on detected clusters (objects) rather than raw fluorescence intensity. Each molecular cluster is represented by a coordinate, and the analysis evaluates the spatial relationships between these objects. Clusters may originate from a single fluorescent channel or from multiple channels. In a single-channel analysis, SODA assesses whether objects within the same population are spatially organized relative to one another. In a multi-channel analysis, SODA evaluates whether objects from different molecular populations are spatially associated beyond random expectation.
In the above image, a two-channel example is represented with:
- Channel 0 (magenta) = reference population
- Channel 1 (cyan) = query population
Around each object detected via the clustering step in Channel 0 (Ch0), SODA creates concentric distance rings, called “Search Radii”. SODA then counts how many Channel 1 (Ch1) clusters fall within each ring, defined by the Q vector value. This allows the algorithm to determine whether clusters preferentially occur at specific distances, whether they avoid each other, or whether their distribution is random.
Example:
Ring
| Search Radii distance (nm)
| Neighbor count (Q vector)
|
Ring 1
| 0–10
| 0
|
Ring 2
| 10–20
| 1
|
Ring 3
| 20–30
| 4
|
Ring 4
| 30-40
| 1
|
SODA repeats this process for every object in the dataset. All Q vectors are aggregated to calculate the coupling distance, p values, and other parameters.
In a dense sample, molecules confined within the same cellular compartment can appear clustered even if randomly distributed. To avoid false-positive colocalization, SODA estimates what the spatial organization would look like under random conditions in a dense sample. Specifically, when nearby clusters generate overlapping search regions, the same neighbor could otherwise be counted multiple times. SODA mathematically compensates for overlapping rings, repeated neighbor counting, and local density artefacts. This improves robustness in crowded biological environments.
The Clustering and Co-Organization analysis App can be run via CODI on localization-based datasets. It contains all of the steps required to cluster localizations and perform the SODA analysis in a step-by-step workflow. Specifically:
- Drift Correction
- Localization Filtering
- Temporal Grouping
- Annotations
- Clustering
- SODA
All these steps are explained in detail in the general “How to Guides” section within the ONI Service Desk.
Important considerations before running the analysis:
- Default values for localization filters and clustering are selected as a starting point. These values, however, need to be optimized according to the specific sample and acquisition performed.
- Clustering should be run:
- If multiple channels must be treated as separate channels and clustered separately, the default ”merge selected channels” settings should not be used. This is the most common case when the tool is applied to investigate co-organization between two different molecular populations.
- In the event all clusters from all channels must be treated as a single population, then “merge selected channels” shall be selected. Example: when a single cluster requires multiple biomarkers in order to be detected accurately.
- The clustering tool will output all localization-based clusters found based on a constrained clustering or DBSCAN* algorithm (see the How to Guide – Clustering). An ID and associated properties such as centroids, size, and area, will be assigned to each of the clusters.
Example 1 - SODA applied on beads datasets
In the image below, Tetraspeck beads were acquired on an ONI Nanoimager using the 638 nm laser (Channel 1) and 561 nm laser (channel 0). During the acquisition, the stage was shifted by around 50 nm along the y-axis and 40 nm along the x-axis to create a grid pattern image. The two channels were mapped with a 20 nm error such that they would appear not colocalized. The resulting image on CODI was a grid made of paired magenta (channel 1) and cyan (channel 0) localization clusters.
When running SODA, we first expect to find a discrete nanoscale organization between magenta and cyan clusters with a co-clustering distance of 20 nm. Given the grid pattern, we also expect a second organization between clusters of the same channel (i.e., magenta-to-magenta and cyan-to-cyan) with a co-clustering distance in the range of 40 to 50 nm.
Before running the SODA tool, two parameters need to be optimized: the number of rings and the maximum radius.
Number of rings: Defines the number of distance bins. The number of rings shall be chosen in relation to the cluster density. Higher ring numbers mean increasing spatial resolution but may also increase noise.
Maximum radius: Defines the largest analyzed distance. Choose a radius relevant to the expected biological organization.
By default, these values are set to 10 and 500, respectively. The values need to be updated based on the sample and the localization density. In this data example, a maximum radius of 60 nm with 10 rings was selected through the control widget. Thus, the tool will evaluate clusters with search radii of 6, 12, 18, 24, 30, 36, 42, 48, 54, and 60 nm.
When completed, the tool will provide four output results within the CODI UI as well as three CSV files containing more detailed information. The four CODI UI outputs include:
- Coupling probability: how likely a specific pair of clusters is to represent meaningful spatial association as a function of the Search Radii distance, rather than random proximity.
- Co-clustering distance: the strongest statistically significant coupling.
- Coupling distance: the average distance of all statistically significant couplings.
- Coupling index: a global summary of how strongly two molecular populations are spatially organized together.
See the section “Understanding the SODA Outputs” for more information.
In this example, the outcome is a co-clustering detected at 18 nm, as expected by the distance of the magenta (Ch1) and cyan (Ch0) pairs, which was around 20 nm. A second peak around 40 nm is found due to the secondary organization distance of neighboring clusters. The coupling distance displayed is 26.69 nm indicating a nanoscale organization, with a coupling index of 1.48.
Merged channels
When applying the merged channels settings in the same bead sample as above, but with 20 number of rings and maximum radius of 60 nm, all the clusters in channels 0 and 1 are treated as belonging to the same population (channel0+channel1). Thus, the outcome is different, with three major coupling probability peaks observed at around16 nm, 40 nm and 50 nm.
Example 2 - SODA applied to nuclear pores
SODA was run on nuclear pore complexes (NPCs) imaged using two-color dSTORM, with nucleoporin protein NUP98-GFP (channel 0, cyan) and the membrane nucleoporin NUP210 (channel 1, magenta, anti-NUP210-Alexa Fluor 647) as the primary targets.
From the literature, it is known that NUP98 occupies the central area of a nuclear pore (Chatel et al. 2012) while NUP210 is found in the periphery (Löschberger et al. 2012). It is therefore expected that when imaged, they will look colocalized.
Given the diameter of a nuclear pore was around 150 nm, a maximum radius of 160 nm with 35 rings was selected, giving the following search radii: 5, 10, 15, 20, ... until 160. The reason is that we expect a small co-clustering distance.
From the image above, significant co-clustering at 15 nm with a 0.0001 p-value was found. The average coupling distance reported was ~34 nm with a coupling index of 0.9, suggesting that almost but not all NUP98 had one coupled neighbor with NUP210.
Understanding the SODA Outputs
After a SODA analysis is completed in CODI, three CSV files are generated. These files contain different levels of the statistical analysis used to compute the coupling plots and summary metrics displayed in the UI.
The CSVs represent:
- co-organization_results.csv: the per-cluster neighbors counts at each Search Radii
- co-organization_vectors.csv: distance-dependent statistical vectors
- Co-organization_summary.csv: final global summary of the analysis
Together, they describe how molecular clusters are spatially organized relative to each other across increasing distances.
co-organization_results.csv
The co-organization_results.csv file contains the per-cluster measurements used internally by the SODA analysis. Each row corresponds to one detected cluster, and the file stores both the cluster properties, and the neighbor counts used to compute the coupling statistics shown in the CODI UI.
The first columns describe the detected cluster itself. Depending on the upstream clustering pipeline, these may include cluster coordinates, cluster area, density, circularity, radius of gyration, and localization statistics.
The key SODA-related columns are the Q vector columns. These columns represent the number of neighboring clusters detected within each search radius around the current cluster for both channel 0 and channel 1. For example, a Q vector of 1 at 36 nm means that one neighboring cluster from the target channel was detected within the corresponding distance ring around that reference cluster.
co-organization_vectors.csv
This file contains the radial statistical vectors used to generate the coupling plot. Each row corresponds to one search radius (distance ring, called “Search Radii”). If:
- maximum radius = 200 nm,
- number of rings = 20,
then each ring represents a 10 nm interval. The Search Radii is used in the x-axis of the CODI coupling plot.
The Raw Q Vector: contains the raw neighbor counts detected at each distance ring. It represents how many neighboring clusters are detected around all reference clusters before statistical correction. Higher values indicate more neighbors detected at that distance. Because larger rings naturally contain more area and therefore more neighbors, these raw values are not directly biologically interpretable on their own.
The p Vector: contains the statistical weighting associated with each distance ring. It represents how strongly the observed neighbor enrichment differs from random expectation. Large values indicate statistically enriched spatial organization at that distance.
The coupling distance displayed in CODI is derived by the above parameters.
Coupling distance: the average distance of all statistically significant coupling. It is defined as
Where,
MDPR = “Mean Distance per Ring” and is defined as the sum of all pairwise distances (across the entire region of interest or field of view) for a given ring, divided by the q_vector.
Coupling weight = is the product between the Raw Q Vector and the p Vector.
The coupling distance can be used to indicate at what spatial scale the overall organization occurs. Example: a small coupling distance of 20-50 nm may indicate nanoscale organization or molecular complexes. A large coupling distance of 300 – 500 nm may indicates compartment organization, mesoscale clustering.
Important: Coupling does not mean binding. SODA identifies statistically enriched spatial associations or events, it does not directly demonstrate molecular binding, physical contact, or biochemical interaction.
Normalized Q Vector: represents the corrected spatial enrichment signal after accounting for density, geometry (such as ring area) and overlap effects. This is the main distance-dependent coupling signal used by the analysis.
The normalized Q vector can be plotted outside of CODI using the csv output and used to better understand the sample, here are some examples:
Coupling profile
| Co-Organization Type
| Interpretation
|
Flat profile near zero
| None
| The neighbor distribution is similar to random distribution. No strong spatial organization is detected.
|
Sharp positive peak at short distances (for example 0–20 nm)
| Nanoscale
| Strong local nanoscale association. Objects are repeatedly found close together more often than expected by chance.
|
Positive peak at larger distances
| Mesoscale
| Preferred spatial separation between object populations. May indicate compartment-level or mesoscale organization.
|
Broad positive plateau across multiple distances
| Diffuse
| Diffuse or extended spatial association over a larger spatial range.
|
Negative dip at short distances
| Exclusion
| Spatial exclusion or steric separation. Fewer neighboring objects are detected than expected at a given distance. Objects systematically avoid being too close or physically cannot occupy the same space.
|
Multiple positive peaks
| Periodic
| Multiple characteristic spatial scales are present. This may occur in periodic, hierarchical, or lattice-like structures.
|
Important: While in the CSV file, both positive and negative values will be reported, the CODI UI shows only the positive values.
Important:
- The co-clustering distance, displayed in CODI, corresponds to the Search Radii value with the highest p-vector value, represents how strongly the observed neighbor enrichment differs from random expectation.
- The coupling probability, displayed in the CODI plot, corresponds to the Normalized Q vectors values after the non-zero Donoho–Johnstone threshold (Donoho & Johnstone 1994) is applied as defined below,
to compensate for the probability of observing random peaks across the different rings.
co-organization_summary.csv
This file contains the high-level summary shown in the CODI UI. Each row corresponds to one channel-to-channel comparison. For example, Channel 0 to Channel 1 means clusters from Channel 0 were used as reference centers, while clusters from Channel 1 were counted around them. The key outputs are:
- Estimated Coupling Distance: represents the average distance between statistically coupled cluster pairs.
Coupling Index: summarizes the overall strength of spatial organization between the two populations
where:
- the q_vector contains neighbor counts at each distance ring,
- the p_vector weights how statistically significant each ring is.
- Number of center points is the total reference population clusters number.
Conceptually, the coupling index represents the average number of statistically significant neighbor relationships per reference cluster.
Coupling index value
| Interpretation
|
0 | No significant coupling
|
<1 | Fewer than one coupled neighbor per object on average
|
1
| Approximately one coupled neighbor per object
|
>1
| Objects significantly coupled to multiple neighbors.
|
p-Value: This measures the statistical significance of the detected organization. Very small values indicate that the observed spatial arrangement is highly unlikely to occur from random distribution alone.
p-Value
| Interpretation
|
> 0.05
| No statistically significant spatial organization detected
|
< 0.05
| Statistically significant spatial organization
|
<0.01
| Strong evidence of spatial organization
|
<0.001
| Very strong evidence that the observed organization is not random
|
References
- Lagache T et al. Mapping molecular assemblies with fluorescence microscopy and object-based spatial statistics. Nat Commun 9, 698 (2018). doi: 10.1038/s41467-018-03053-x
- Donoho DL & Johnstone IM. Ideal spatial adaptation by wavelet shrinkage. Biometrika, 81(3): 425-455 (1994).
- Chatel G et al. Domain topology of nucleoporin Nup98 within the nuclear pore complex. Journal of Structural Biology (2012).
- Löschberger A. et al. Super-resolution imaging visualizes the eightfold symmetry of gp210 proteins around the nuclear pore complex and resolves the central channel with nanometer resolution. J. Cell Sci. 125, 570–575 (2012)