In progress

Evaluation project

GeoEval Benchmark

A benchmark design for testing whether AI systems can execute realistic GIS tasks with reproducible, reviewable results.

Current scope

GeoEval is an in-progress benchmark concept for evaluating geospatial AI against real GIS work. Public project material will exclude confidential client data, proprietary datasets, and unsupported performance claims.

In-progress project

This page describes work in development. It does not claim published benchmark results, client work, or product performance.

Planned deliverables

What this work is designed to produce.

  • 01

    Task library

  • 02

    Inputs and reference outputs

  • 03

    Scoring rubric

  • 04

    Failure taxonomy and regression pack

Discuss an Evaluation Pilot