Addressing the bottleneck of data overload
Article

Addressing the bottleneck of data overload

Crowdsourced seafloor mapping: federated learning at the edge

Crowdsourced bathymetry is increasingly seen as a way to expand seabed coverage by collecting depth measurements from vessels during routine operations. Yet as data volumes grow, the primary challenge shifts from acquisition to validation: how can large amounts of opportunistic, heterogeneous measurements be processed efficiently and reliably? Drawing on deployments in Danish waters, this article explores how edge computing and federated learning can address this bottleneck, enabling automated anomaly detection while allowing hydrographic experts to focus on the most relevant observations for seabed monitoring and survey prioritization.

According to the latest Seabed 2030 update, only 27.3% of the seabed has been mapped to modern standards. With dedicated vessels operating along dense, controlled track lines, today’s hydrographic surveys produce high-quality data, but also at high cost and on a limited scale. As a result, survey efforts focus on high-priority areas for navigational safety, leaving large regions unmapped or covered by older surveys. Further complicating matters, the seabed is far from static. Currents and waves continually reshape sandy or soft bottoms, while human activities – such as port expansions, channel dredging and the construction of offshore wind farms – introduce further changes to the marine environment.

Crowdsourcing to increase capacity

Addressing the limitations of current survey capacity cannot be achieved by dedicated survey vessels alone. Instead, broader participation and innovation are required. Therefore, for Danish waters, where nearly three quarters of the seabed remain unmapped to modern standards, the Danish Geodata Agency (DGA) has been exploring the possibility of leveraging crowdsourced bathymetry. This approach, in which ordinary vessels contribute depth measurements during routine operations, offers the potential to increase coverage, keep pace with dynamic seabed changes and gradually validate older datasets. It can also serve as a decision-support tool to prioritize resurveying efforts and, where data quality meets required standards (e.g. IHO S-44), contribute to official charting.   

To address these challenges, the DGA joined the Horizon Europe MobiSpaces project, which developed data governance, analytics and edge computing solutions across mobility domains. In the ‘CrowdSeaMapping’ use case, the DGA and the Austrian Institute of Technology (AIT) explored how crowdsourced depth data can be integrated into workflows using machine learning to detect errors and anomalies.

In the geospatial domain, projects such as OpenStreetMap have mapped large parts of the world through voluntary contributions – in some cases surpassing the detail available in official national or commercial datasets. A comparable approach in the marine sector is defined by the International Hydrographic Organization (IHO) as Crowdsourced Bathymetry (CSB). Most vessels are already equipped with echosounders and GNSS, but the lack of dedicated systems for data capture and transmission remains a key limitation. High-speed 4G or 5G connectivity is largely confined to coastal waters, while satellite bandwidth is costly. Practical solutions therefore rely on onboard storage of raw data with deferred transmission when connectivity becomes available.

Figure 1: The federated learning data pipeline.

Processing at the edge

Crowdsourcing introduces additional challenges. Data collection is unsupervised, and as participation scales, so do data volumes and processing demands, making cleaning labour-intensive. To address this, the DGA and AIT explored a federated learning approach in which processing is performed directly at the data collector. This reduces transmission requirements and shifts data cleaning to the edge nodes of the system (Figure 1). This architecture avoids the need to centralize large volumes of raw sensor data while still allowing model updates to benefit from distributed observations.

Raw CSB data is prone to artefacts such as false bottoms from double returns, spikes from aeration or cavitation, offsets from unmodelled draught or tides, and occasional timing or sensor errors. Normally, manual data cleaning is necessary but this is unfeasible at scale. Therefore, the CrowdSeaMapping approach proposes that data collectors employ an onboard artificial intelligence (AI) model termed ‘MapFed’ which enables continuous modelling of the seafloor as well as automatic detection of anomalies.

Detecting anomalies in real time

MapFed learns the expected depth distribution for each location using an adaptive prototype approach. The model is initialized from an existing bathymetric grid (in this case, the Danish Depth Model) and can be continuously trained with incoming survey and/or CSB data. Measurements that deviate significantly from the learned distribution are flagged locally. Flagged points are then submitted for expert review. Hydrographers assess whether these represent true artefacts to be discarded or valid deviations that should be assimilated into updated bathymetric grids. This workflow creates a feedback loop; domain expertise continuously improves both the anomaly detection model and the underlying bathymetric reference.

Several system components were evaluated both at sea and in controlled environments. The data collector was field-tested aboard a vessel, demonstrating reliable operational performance. Meanwhile, the MapFed anomaly detection model and federated learning setup were validated in the lab, confirming feasibility and potential to reduce data storage and transmission requirements. Together, these trials demonstrate the technical feasibility of the approach and its potential for scalable crowdsourced bathymetry.

Figure 2: Edge device installed aboard R/V Dana IV.

The edge device

The data collector, developed by the Danish company Sternula, interfaces with a vessel’s NMEA network, listening to the onboard sensors such as echosounder and GNSS. It incorporates a Raspberry Pi compute module for running the MapFed model (Figure 2), with geofencing applied to restrict data collection to the Danish Exclusive Economic Zone (EEZ).

Testing at sea

The system was first tested in the summer of 2023 on board the research vessel Dana IV (DTU Aqua). During this trial, the data collector operated continuously and without failure for 37 days (Figure 3). The collected CSB data proved immediately useful in resolving discrepancies between two conflicting surveys of Skagerrak. In addition, the results were later incorporated into the second version of the Danish Depth Model (DDM) which was released in August 2024, demonstrating that cleaned CSB data provided better input than interpolated estimates. While the coverage was limited – approximately 8,500 grid cells (50×50m), corresponding to approximately 21km² – this represented a significant first step for the DGA.

A second, longer deployment took place between April 2024 and March 2025. Over this period, the system again proved robust, operating for nearly one year (300 days at sea) without intervention. During this test, depth data from large areas of the North Sea was collected which has the potential to be used in future versions of the DDM.

Figure 3: The track line of the first sea trial, with the Danish Depth Model v2 as the background.

Comparing AI and expert cleaning

To assess MapFed’s performance, data from the second deployment was manually cleaned by a hydrographic expert and compared with the model output. The results showed strong agreement with expert cleaning (over 85% of records), indicating that most artefacts can be reliably identified without manual intervention. The model was trained on the public DDM, which has 50m resolution and variable data quality – from high-quality multibeam surveys to interpolated estimates based on historical lead-line measurements. To avoid propagating uncertainty, interpolated values were excluded during training, leaving gaps where no reliable reference data existed. As expected, model performance was lowest in areas of low DDM reliability, where limited training data led to misclassifications.

Since no comprehensive ground truth dataset exists for Danish waters, validation relies on expert review, supported by contextual information beyond the public DDM. This feedback not only provides a benchmark for anomaly detection but also supports iterative improvement of both the bathymetric grids and the MapFed model. Over time, these gaps can be reduced through additional CSB contributions, while the MapFed approach enables continuous model improvement as participation increases.

Towards scalable seabed mapping

Initial tests of the data collector and the MapFed model demonstrate that crowdsourced bathymetry is both technically feasible and operationally valuable. Even limited deployments have already shown how CSB can validate existing surveys, improve national bathymetric models and highlight discrepancies that would otherwise remain hidden. Across the two deployments, more than 5.7 million depth points were collected, covering approximately 4.3% of the 50×50m grid cells in the DDM. This demonstrates that, at sufficient scale, CSB can make a measurable contribution to national mapping.

With broader participation from merchant vessels, fishing boats and larger yachts, CSB could provide a continuous stream of depth data that keeps pace with the dynamic seabed. Combined with federated learning and edge processing, this creates a sustainable model for large-scale, data integration. Rather than replacing national hydrographic surveys, CSB complements them, supporting prioritization of resurveys. In the future, if quality thresholds are met, it could also serve as an input to nautical charts. Therefore, with sufficient adoption and coordination, CSB has the potential to significantly accelerate progress toward global initiatives such as Seabed 2030, while also giving hydrographic offices a practical way to keep bathymetric reference data up to date. It represents not just a method to close mapping gaps, but a shift toward a more dynamic, participatory and data-rich approach to understanding and managing the marine environment.

Figure 4: Track line from second sea trial in 2024-2025, with the Danish Depth Model V2 as the background.

Further reading

Graser, A., Heistracher, C., & Pruckovskaja, V. (2022). On the Role of Spatial Data Science for Federated Learning, Spatial Data Science Symposium (SDSS2022). https://doi.org/10.25436/E24K5T

MobiSpaces – New data spaces for green mobility. https://cordis.europa.eu/project/id/101070279

Graser, A., Weißenfeld, A., Heistracher, C., Dragaschnig, M., & Widhalm, P. (2024) Federated Learning for Anomaly Detection in Maritime Movement Data. 25th IEEE International Conference on Mobile Data Management (MDM2024), 24-27 June 2024, Brussels, Belgium. doi:10.1109/MDM61037.2024.00030

Graser. A., Widhalm, P., & Dragaschnig, M. (2020). The M³ massive movement model: a distributed incrementally updatable solution for big movement data exploration, International Journal of Geographical Information Science, 34(12), 2517-2540. doi:10.1080/13658816.2020.1776293

Masetti, Giuseppe & Rondeau, Mathieu & Barón, Belén & Wills, Peter & Petersen, Yvonne & Salmia, Juho. (2020), Trusted Crowd-Sourced Bathymetry: From the Trusted Crowd to the Chart. 10.13140/RG.2.2.36642.86722.

Acknowledgements

This work was carried out as part of the EU Horizon Europe project, MobiSpaces (Grant Agreement No. 101070279).

Hydrography Newsletter

Value staying current with hydrography?

Stay on the map with our expertly curated newsletters.

We provide educational insights, industry updates, and inspiring stories from the world of hydrography to help you learn, grow, and navigate your field with confidence. Don't miss out - subscribe today and ensure you're always informed, educated, and inspired by the latest in hydrographic technology and research.

Choose your newsletter(s)