---
title: Core-S2L2A-249k-SatCLIP
canonical_url: "https://www.modelscope.cn/datasets/Major-TOM/Core-S2L2A-249k-SatCLIP"
md_url: "https://www.modelscope.cn/datasets/Major-TOM/Core-S2L2A-249k-SatCLIP.md"
repository: Major-TOM/Core-S2L2A-249k-SatCLIP
last_updated: 2026-05-23
license: cc-by-sa-4.0
storage_size: "300 MB"
downloads: 166
stars: 1
---

# Core-S2L2A-249k-SatCLIP

> Core-S2L2A-249k-SatCLIP - Major-TOM 在 ModelScope 开源的数据集。Core-S2RGB-249k-SatCLIP

Major-TOM/Core-S2L2A-249k-SatCLIP 是 ModelScope 魔搭社区上的数据集，存储大小 300 MB，采用 cc-by-sa-4.0 许可。

- **Repository**: Major-TOM/Core-S2L2A-249k-SatCLIP
- **License**: cc-by-sa-4.0
- **Storage size**: 300 MB
- **Downloads**: 166
- **Stars**: 1
- **Last updated**: 2026-05-23

Source: https://www.modelscope.cn/datasets/Major-TOM/Core-S2L2A-249k-SatCLIP

---

# Core-S2RGB-249k-SatCLIP

Geospatial-vision embedding dataset computed from **Core-S2L2A-249k** using the **SatCLIP** model.

## Overview

| Property | Value |
|----------|-------|
| Source imagery | Core-S2L2A-249k (248,719 patches, 384 × 384 px) |
| Model | SatCLIP (ResNet-50 + location encoder) |
| Input bands | All 12 Sentinel-2 bands [B01, B02, B03, B04, B05, B06, B07, B08, B8A, B09, B11, B12] |
| Embedding dimension | 512 |
| Output format | GeoParquet |
| License | CC-BY-SA-4.0 |

## Computation Pipeline

1. **Pre-processing**: Each 384 × 384 Sentinel-2 L2A patch is read from the source parquet files. All 12 spectral bands are stacked and normalised by dividing by `1e4` to convert digital numbers to reflectance.
2. **Resize & Pad**: The 12-band tensor is interpolated to the SatCLIP input size of **224 × 224** pixels. Because SatCLIP expects 13 input channels, a zero-filled B10 channel is padded at index 10.
3. **Encoding**: The 13-channel tensor is fed into the SatCLIP image encoder (a ResNet-50 trained with location-aware contrastive learning on satellite imagery) to extract a 512-dimensional image embedding.
4. **Post-processing**: No L2-normalisation is applied during dataset generation; normalisation is performed at retrieval time if required.
5. **Geospatial metadata**: The original UTM footprint is reprojected to EPSG:4326 (WGS-84) to obtain the `geometry`, `centre_lat`, and `centre_lon` fields. Additional metadata (`product_id`, `grid_cell`, `timestamp`, `utm_crs`, `pixel_bbox`) is preserved from the source dataset.

## File Layout

```
.
├── SatCLIP_crop_384x384.parquet   # Main embedding GeoParquet (248,719 rows)
└── README.md
```

## Schema

| Column | Type | Description |
|--------|------|-------------|
| `unique_id` | string | SHA-256 hash of geometry + timestamp + product_id + embedding |
| `embedding` | list<float> | 512-dim SatCLIP feature vector |
| `timestamp` | datetime | Acquisition timestamp |
| `product_id` | string | Original Sentinel-2 product identifier |
| `grid_cell` | string | Major-TOM grid cell identifier |
| `grid_row_u` | int16 | Grid row index |
| `grid_col_r` | int16 | Grid column index |
| `geometry` | geometry | WGS-84 polygon (footprint) |
| `centre_lat` | float32 | Latitude of patch centre |
| `centre_lon` | float32 | Longitude of patch centre |
| `utm_footprint` | string | Original UTM footprint as WKT |
| `utm_crs` | string | UTM CRS (e.g. EPSG:32633) |
| `pixel_bbox` | list<int> | Pixel bounding box [x_min, y_min, x_max, y_max] |
| `parquet_url` | string | Source parquet file path in the image dataset |
| `parquet_row` | int64 | Row index within the source parquet file |

## Usage

```python
import pandas as pd

df = pd.read_parquet("SatCLIP_crop_384x384.parquet")
print(len(df), "embeddings")
print(df.iloc[0].embedding.shape)  # (512,)
```

## Citation

If you use this embedding dataset, please cite the original Major-TOM paper and the SatCLIP paper:

```bibtex
@article{zheng2026earthembeddingexplorer,
  title={EarthEmbeddingExplorer: A Web Application for Cross-Modal Retrieval of Global Satellite Images},
  author={Zheng, Yijie and Wu, Weijie and Wu, Bingyue and Zhao, Long and Li, Guoqing and Czerkawski, Mikolaj and Klemmer, Konstantin},
  journal={arXiv preprint arXiv:2603.29441},
  year={2026},
  note={ICLR 2026 Workshop ML4RS Tutorial Track (oral)}
}
```

```bibtex
@inproceedings{francis2024majortom,
  title={Major TOM: Expandable Datasets for Earth Observation},
  author={Francis, Alistair and Czerkawski, Mikolaj},
  year={2024},
  booktitle={IGARSS 2024},
  eprint={2402.12095},
  archivePrefix={arXiv}
}
```
