- 人工智能
- 计算机视觉
- GIS
- 图像处理
- 微调
【免费下载链接】geoai
GeoAI: Artificial Intelligence for Geospatial Data
geoai.auto是 GeoAI 中基于 Hugging Face 基础模型(transformers)构建的自动推理模块,它通过AutoGeoModel与AutoGeoImageProcessor两个核心类,把「模型自动加载 + GeoTIFF 地理元数据读写」封装成一套统一 API,让语义分割、目标检测、图像分类、深度估计与掩码生成等任务无需人工标注或训练即可直接作用于遥感影像。读完本文,你将掌握如何用几行代码完成零样本目标检测(如 Grounding DINO)、语义分割(如 SegFormer)、图像分类(如 ViT)与深度估计,并把栅格结果原地矢量化导出为 GeoJSON 等地理矢量格式。
一、auto 模块的定位与设计目标
按照官方文档 docs/auto.md 的说明,auto模块提供基于 Hugging Face 基础模型的自动分割(auto-segmentation)能力:它利用预训练模型对地理空间影像执行自动图像分割,无需手工提示词(manual prompts)或训练数据,面向零样本(zero-shot)与小样本(few-shot)分割任务开箱即用。
模块级 docstring(见 geoai/auto.py 开头)进一步明确了实现思路:AutoGeoModel与AutoGeoImageProcessor分别扩展 transformers 的AutoModel与AutoImageProcessor,为其增加:
- 处理 GeoTIFF 等栅格数据的能力(保留 CRS、仿射变换、边界等地理信息);
- 将推理结果写回GeoTIFF 栅格或矢量数据(GeoJSON、GPKG、Shapefile 等)的能力。
该模块常与示例笔记本 docs/examples/AutoModel.ipynb 搭配使用,后者给出了从下载样例影像到完成三类任务、再到结果落盘的完整可运行流程。
二、支持的模型任务与 TASK_MODEL_MAPPING
AutoGeoModel内部通过TASK_MODEL_MAPPING(geoai/auto.py 中TASK_MODEL_MAPPING)把任务名映射到对应的 transformers Auto 类:
| 任务名(task 参数) | 对应 transformers 类 | 典型模型 |
|---|---|---|
segmentation/semantic-segmentation | AutoModelForSemanticSegmentation | SegFormer、Mask2Former |
image-segmentation | AutoModelForImageSegmentation | 通用图像分割模型 |
universal-segmentation | AutoModelForUniversalSegmentation | Mask2Former 等统一分割模型 |
depth-estimation | AutoModelForDepthEstimation | Depth Anything、DPT |
mask-generation | AutoModelForMaskGeneration | SAM(Segment Anything) |
object-detection | AutoModelForObjectDetection | DETR、YOLOS |
zero-shot-object-detection | AutoModelForZeroShotObjectDetection | Grounding DINO、OWL-ViT |
classification/image-classification | AutoModelForImageClassification | ViT、ResNet |
如果你在from_pretrained中不指定task,模块会尝试从 Hugging Face 模型配置(AutoConfig)的architectures字段自动推断:包含Segmentation字样走语义分割,包含Detection走目标检测,包含Classification走图像分类,否则回退到AutoModel。这一推断逻辑位于_load_model_from_pretrained(geoai/auto.py 中_load_model_from_pretrained),并通过tests/test_auto.py的test_task_model_mapping用例验证了各任务键的完整性。
三、AutoGeoImageProcessor:带地理元数据的图像处理器
AutoGeoImageProcessor包装 Hugging Face 的AutoImageProcessor(或需要文本输入时的AutoProcessor),负责「读取栅格 → 预处理 → 交给模型 → 写回 GeoTIFF」全链路的像素与元数据管理。它遵循 transformers 惯例,通过from_pretrained实例化:
from geoai.auto import AutoGeoImageProcessor processor = AutoGeoImageProcessor.from_pretrained( "nvidia/segformer-b0-finetuned-ade-512-512", device="cuda", # 不传则自动探测(见 geoai.utils.get_device) )关键设计要点:
- 处理器自动选择:
from_pretrained的use_full_processor=True会强制使用AutoProcessor(同时接收文本与图像);当模型名包含grounding-dino、owl、clip、blip等关键字时也会自动启用,因为这类模型需要文本输入。若AutoImageProcessor加载失败,会静默回退到AutoProcessor。 - load_geotiff:接收 GeoTIFF 路径或已打开的
rasterio.DatasetReader,可传window(rasterioWindow)只读子窗口、bands(1 起始的波段索引列表)选择波段;返回(CHW 数组, 元数据字典),元数据含profile、crs、transform、bounds、width、height、count。 - load_image:统一入口,自动区分 GeoTIFF(按 CRS 或
.tif/.tiff后缀判定)与普通图片(PIL 加载并转 RGB);也接受numpy.ndarray(HWC 或 2D 自动规整为 CHW)与 PILImage,不支持的输入类型会抛出TypeError。 - prepare_for_model:预处理阶段会把 CHW 转回 HWC,单波段重复三次、超过 3 波段截断;默认开启
normalize=True且采用百分位裁剪(对每个波段取 2% 与 98% 分位数做 min-max 拉伸到 [0, 255]),对遥感影像常见的低对比度、饱和波段非常友好;对于 p2 == p98 的常量波段会安全置零(tests/test_auto.py的test_prepare_percentile_clip_uniform_band专门覆盖此边界)。 - save_geotiff:把推理结果连同原始
metadata["profile"]写回 GeoTIFF,支持dtype、compress(默认"lzw")、nodata参数,并自动创建输出目录。2D 数组作为单波段写入,3D CHW 数组按多波段写入。
tests/test_auto.py对上述能力均有对应测试:test_load_geotiff_from_path、test_load_geotiff_with_bands、test_load_geotiff_with_window、test_load_image_from_png、test_prepare_for_model_4band_truncates、test_save_geotiff_3d等。
四、AutoGeoModel:核心推理类与 predict 全参数
AutoGeoModel是模块的主入口,from_pretrained加载模型后即自动移到设备并置为eval()模式:
from geoai.auto import AutoGeoModel model = AutoGeoModel.from_pretrained( "IDEA-Research/grounding-dino-base", task="zero-shot-object-detection", device="cuda", # 可选 tile_size=1024, # 大图分块尺寸,默认 1024 overlap=128, # 分块重叠像素,默认 128 ) result = model.predict("aerial.tif", text="a building. a car.", box_threshold=0.3)predict 方法参数一览
| 参数 | 默认值 | 说明 |
|---|---|---|
source | — | 输入:文件路径(含 http/https URL)、numpy 数组或 PIL Image |
output_path | None | 输出 GeoTIFF 路径(需输入带地理元数据才会写栅格) |
output_vector_path | None | 输出矢量路径(GeoJSON、GPKG 等),自动把掩码/检测框转矢量 |
window/bands | None | rasterio 子窗口与波段选择 |
threshold | 0.5 | 分割掩码二值化阈值 |
text | None | 零样本检测文本提示,如"a building. a tree."(小写且以句号结尾) |
labels | None | 标签列表,传入后自动转换为 text 格式 |
box_threshold | 0.3 | 检测框置信度阈值 |
text_threshold | 0.25 | 零样本检测文本相似度阈值 |
min_object_area | 100 | 矢量化时保留的最小对象面积(像素) |
simplify_tolerance | 1.0 | 多边形简化容差(像元单位,乘以像元尺寸) |
batch_size | 1 | 分块推理批大小 |
return_probabilities | False | 是否在结果中附带概率图 |
内部处理路径分为三类(见 geoai/auto.py 中predict与_predict_detection):
- 零样本 / 常规目标检测:走
_predict_detection,使用processor.post_process_grounded_object_detection(Grounding DINO 类模型)或post_process_object_detection(DETR 类模型)解码,得到boxes、scores、labels;若模型不支持 grounded 后处理,代码会自动降级为基于pred_boxes+ logits sigmoid 的兜底方案。检测框可通过output_vector_path导出为 GeoDataFrame(坐标在无地理元数据时为像素空间)。 - 分块推理:当影像尺寸超过
tile_size(默认 1024)且任务不是分类时,自动启用_predict_tiled分块处理;每个块的有效尺寸为tile_size - 2 * overlap,块与块之间重叠overlap像素,重叠区通过「累加再求平均」平滑融合,避免边缘伪影(seam artifacts);配合 tqdm 进度条逐块推进,块失败仅告警不中断。 - 单图推理:尺寸未超限的影像走
_predict_single,在torch.no_grad()下前向一次得到结果。
_process_outputs负责把不同模型的原始输出归一化为统一字典:4Dlogits→argmax得到语义分割掩码mask;2Dlogits→ 分类class与probabilities;pred_masks→ 掩码;predicted_depth→depth;SAM 风格的masks→ 二值掩码与all_masks。tests/test_auto.py的test_process_outputs_segmentation_logits、test_process_outputs_classification_logits、test_process_outputs_depth分别验证了这三条分支。
掩码矢量化:mask_to_vector
mask_to_vector(geoai/auto.py 中mask_to_vector)是「栅格结果转矢量」的关键:它基于rasterio.features.shapes提取连通域多边形,按min_object_area/max_object_area(像素面积)过滤碎小对象与超大连通域,再按simplify_tolerance(乘以像元尺寸、preserve_topology=True)简化边界,最终返回带 CRS 的 GeoDataFrame。若缺少 CRS 或 transform、掩码为空,会分别返回None并给出日志警告,对应tests/test_auto.py中test_mask_to_vector_no_crs、test_mask_to_vector_empty_mask等用例。save_vector则按扩展名自动选择驱动:.geojson/.json→ GeoJSON,.gpkg→ GPKG,.shp→ ESRI Shapefile,.parquet→ Parquet,.fgb→ FlatGeobuf,缺省回退 GeoJSON。
五、四个便捷函数:一行代码完成典型任务
除了底层类,模块还导出了四个函数式入口(签名与默认值均以 geoai/auto.py 为准):
1. semantic_segmentation —— 语义分割
from geoai.auto import semantic_segmentation result = semantic_segmentation( "aerial.tif", output_path="segmentation_output.tif", model_name="nvidia/segformer-b0-finetuned-ade-512-512", # 默认模型 output_vector_path="segmentation.geojson", # 可选:同步矢量化 threshold=0.5, tile_size=1024, overlap=128, min_object_area=100, simplify_tolerance=1.0, device=None, # 自动探测 cuda/cpu ) mask = result["mask"] # (H, W) uint8 掩码2. depth_estimation —— 深度估计
from geoai.auto import depth_estimation result = depth_estimation( "aerial.tif", output_path="depth_output.tif", model_name="depth-anything/Depth-Anything-V2-Small-hf", # 默认模型 ) depth = result["depth"] # 相对深度图3. image_classification —— 图像分类
from geoai.auto import image_classification result = image_classification( "aerial.tif", model_name="google/vit-base-patch16-224", # 默认模型 ) cls_index = result["class"] # 预测类别索引 probs = result["probabilities"] # 各类别概率向量注意:分类任务不会触发分块推理(predict中对 classification 任务做了显式排除),因此适合整幅影像语义判读类场景。
4. object_detection ——(零样本)目标检测
from geoai.auto import object_detection result = object_detection( "aerial.tif", labels=["building", "tree", "car", "road"], # 自动转为 "a building. a tree. ..." model_name="IDEA-Research/grounding-dino-base", # 默认模型 output_vector_path="detections.geojson", # 可选 box_threshold=0.3, text_threshold=0.25, ) boxes, scores, labels = result["boxes"], result["scores"], result["labels"]object_detection内部会根据模型名自动选择任务:模型名含grounding-dino或owl时走zero-shot-object-detection,否则走object-detection。labels列表会被自动拼接为以小写、句号结尾的提示语格式。
六、完整实战流程(继承自 AutoModel 示例)
以下流程对应官方示例 docs/examples/AutoModel.ipynb,可完整落地运行:
# 1. 下载样例航空影像并查看 from geoai import download_file image_url = "https://huggingface.co/datasets/giswqs/geospatial/resolve/main/aerial.tif" image_path = download_file(image_url, "aerial.tif") # 2. 零样本目标检测 + 矢量导出 result = object_detection( image_path, labels=["building", "tree", "car", "road"], box_threshold=0.25, text_threshold=0.25, output_vector_path="detections.geojson", ) print(f"Detected {len(result.get('boxes', []))} objects") # 3. 语义分割 + 栅格/矢量双输出 seg_result = semantic_segmentation( image_path, output_path="segmentation_output.tif", output_vector_path="segmentation.geojson", model_name="nvidia/segformer-b0-finetuned-ade-512-512", min_object_area=50, simplify_tolerance=1.0, ) print(f"Mask shape: {seg_result['mask'].shape}, " f"polygons: {len(seg_result['geodataframe'])}") # 4. 图像分类 + Top-5 解释 import numpy as np from transformers import AutoConfig cls_result = image_classification(image_path, model_name="google/vit-base-patch16-224") config = AutoConfig.from_pretrained("google/vit-base-patch16-224") probs = cls_result["probabilities"] top5 = np.argsort(probs)[-5:][::-1] for idx in top5: print(f"{config.id2label.get(idx, f'Class {idx}')}: {probs[idx]:.4f}")七、结果可视化辅助函数
auto模块附带四组基于 matplotlib 的展示函数,便于快速目检结果(实现见 geoai/auto.py 可视化部分):
show_image(source, figsize=(10, 10), title=None, axis_off=True):显示 GeoTIFF 或普通图片,GeoTIFF 会自动做百分位拉伸转 uint8;show_detections(source, detections, ...):在原图上叠加检测框与标签/分数,支持box_color单色或列表配色、show_scores开关;show_segmentation(source, mask, ...):原图与掩码叠加并排显示(show_original=True时左右分栏),掩码尺寸不一致时自动用最近邻插值对齐;show_depth(source, depth, ...):深度图用plasma色带渲染并附色条。
from geoai.auto import show_detections, show_segmentation show_detections(image_path, result, title="Zero-Shot Object Detection") show_segmentation(image_path, seg_result["mask"], title="Semantic Segmentation", alpha=0.6)八、辅助工具函数
get_hf_tasks():从transformers.pipelines.SUPPORTED_TASKS列出当前 transformers 支持的全部任务名(含未在本模块映射表中的任务,供参考与扩展);get_hf_model_config(model_id):基于AutoConfig.from_pretrained返回模型的完整配置字典,可用于查看model_type、num_labels、hidden_sizes等元信息,例如:
from geoai.auto import get_hf_model_config config = get_hf_model_config("nvidia/segformer-b0-finetuned-ade-512-512") print(config.get("model_type"), config.get("num_labels"))九、源码、测试与生态印证
- 实现主体:geoai/auto.py(约 2000 行),模块顶部通过
hf_logging.set_verbosity_error()压低 transformers 加载日志,默认关闭 HF 加载报告; - 测试覆盖:tests/test_auto.py 覆盖处理器 I/O、预处理边界、输出解析、掩码矢量化、检测转 GeoDataFrame、便捷函数签名与模块导出(
__all__)等约 40 个用例,可作为理解各方法契约的权威参考; - 生态关联:
mask-generation任务可搭配 geoai/segment.py 中的 GroundedSAM 使用文本提示做实例级分割;而 geoai/onnx.py 中的ONNXGeoModel明确以AutoGeoModel的 API 为镜像,两者共用predict语义,方便把同一套工作流切换到 ONNX 推理; - 入口方式:
geoai.auto既可通过from geoai.auto import AutoGeoModel直接导入,也可作为geoai顶层延迟加载符号体系的一部分使用(见 geoai/init.py 的_LAZY_SYMBOL_MAP与_LAZY_SUBMODULES机制,import geoai不会立即拉起 torch/transformers 等重型依赖)。
十、使用前提与注意事项
- 依赖:使用
auto模块需要torch、transformers、rasterio、geopandas、numpy、PIL等依赖已安装,模型权重首次加载时从 Hugging Face Hub 下载,请确保网络可达; - 设备:不显式传
device时由geoai.utils.get_device自动探测(有 CUDA 用 cuda,否则 CPU); - 零样本检测提示语:Grounding DINO 类模型对提示语敏感,建议使用小写、以句号结尾的短语(如
"a building. a car."),并依据影像分辨率调整box_threshold与text_threshold(示例中常用 0.25); - 大图分块:超过
tile_size的影像会自动分块,分块推理的内存峰值取决于单块尺寸,可适当调小tile_size或增大overlap换取更平滑的拼接结果; - 矢量化前提:
mask_to_vector需要输入具备 CRS 与仿射变换(即 GeoTIFF 输入),普通图片由于缺少地理元数据只能停留在像素坐标。
综上,geoai.auto把「模型选择、影像预处理、分块推理、地理化输出、结果可视化」整合为一条完整链路,是 GeoAI 中对 Hugging Face 生态做地理空间适配的核心模块之一,适合快速原型验证、零样本地物提取与自动化遥感批处理任务。
- 人工智能
- 计算机视觉
- GIS
- 图像处理
- 微调
【免费下载链接】geoai
GeoAI: Artificial Intelligence for Geospatial Data
相关推荐
GeoAI MCP Server 实战指南:用 Claude Desktop 等 AI Agent 驱动 GeoAI 地理空间 AI 能力
GeoAI MCP Server 实战指南:用 Claude Desktop 等 AI Agent 驱动 GeoAI 地理空间 AI 能力 GeoAI MCP
人工智能计算机视觉GIS图像处理微调GeoAI ONNX 模块实战指南:PyTorch 模型导出与 GeoTIFF 加速推理
GeoAI ONNX 模块实战指南:PyTorch 模型导出与 GeoTIFF 加速推理 GeoAI 的 ONNX 模块( docs/onnx.md https
人工智能计算机视觉GIS图像处理微调cli-anything-geoai 使用指南:用命令行与 AI Agent 驱动 GeoAI 地理空间分析
cli anything geoai 使用指南:用命令行与 AI Agent 驱动 GeoAI 地理空间分析 cli anything geoai 是 GeoA
人工智能计算机视觉GIS图像处理微调
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考