# Machine Perception

Published articles for Machine Perception.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Teaching AI to read a map

DevFeed: [Teaching AI to read a map](<https://devfeed.tech/articles/teaching-ai-to-read-a-map-6884.md>)

Original publisher: [Read original article](<https://research.google/blog/teaching-ai-to-read-a-map/>)

Published: 2026-02-17T21:37:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Machine Perception](<https://devfeed.tech/topics/machine-perception.md>), [Synthetic Data Generation](<https://devfeed.tech/topics/synthetic-data-generation.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [large-language-models](<https://devfeed.tech/topics/large-language-models.md>), [Google](<https://devfeed.tech/topics/google.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [generation](<https://devfeed.tech/tags/generation.md>), [google](<https://devfeed.tech/tags/google.md>), [images](<https://devfeed.tech/tags/images.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [machine-perception](<https://devfeed.tech/tags/machine-perception.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [open-source-models-datasets](<https://devfeed.tech/tags/open-source-models-datasets.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>)

### AI overview

Google Research presents MapTrace, a task, dataset, and synthetic data generation pipeline for teaching multimodal large language models to trace valid routes on maps. The work addresses models' difficulty with spatial, geometric, and topological reasoning and releases 2 million generated question-answer pairs using Gemini 2.5 Pro and Imagen-4 Models.

### Source excerpt

Machine Perception

## Small models, big results: Achieving superior intent extraction through decomposition

DevFeed: [Small models, big results: Achieving superior intent extraction through decomposition](<https://devfeed.tech/articles/small-models-big-results-achieving-superior-intent-extraction-through-decomposition-6874.md>)

Original publisher: [Read original article](<https://research.google/blog/small-models-big-results-achieving-superior-intent-extraction-through-decomposition/>)

Published: 2026-01-22T16:56:44Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Google](<https://devfeed.tech/topics/google.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [generative](<https://devfeed.tech/tags/generative.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llms](<https://devfeed.tech/tags/llms.md>), [machine-perception](<https://devfeed.tech/tags/machine-perception.md>), [mobile](<https://devfeed.tech/tags/mobile.md>), [mobile-systems](<https://devfeed.tech/tags/mobile-systems.md>), [models](<https://devfeed.tech/tags/models.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [on-device](<https://devfeed.tech/tags/on-device.md>), [performance](<https://devfeed.tech/tags/performance.md>)

### AI overview

Google researchers describe a decomposed approach for extracting user intent from web and mobile UI interaction trajectories with small multimodal language models. The method summarizes individual screens first, then infers overall intent from the sequence of summaries, achieving results comparable to much larger models while supporting on-device processing.

### Source excerpt

Generative AI

## Google Earth AI: Unlocking geospatial insights with foundation models and cross-modal reasoning

DevFeed: [Google Earth AI: Unlocking geospatial insights with foundation models and cross-modal reasoning](<https://devfeed.tech/articles/google-earth-ai-unlocking-geospatial-insights-with-foundation-models-and-cross-modal-reasoning-6797.md>)

Original publisher: [Read original article](<https://research.google/blog/google-earth-ai-unlocking-geospatial-insights-with-foundation-models-and-cross-modal-reasoning/>)

Published: 2025-10-23T08:08:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [Earth AI](<https://devfeed.tech/topics/earth-ai.md>), [Google](<https://devfeed.tech/topics/google.md>), [foundation-models](<https://devfeed.tech/topics/foundation-models.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [object-detection](<https://devfeed.tech/topics/object-detection.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [climate-sustainability](<https://devfeed.tech/tags/climate-sustainability.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [developers](<https://devfeed.tech/tags/developers.md>), [earth-ai](<https://devfeed.tech/tags/earth-ai.md>), [foundation-models](<https://devfeed.tech/tags/foundation-models.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [google](<https://devfeed.tech/tags/google.md>), [machine-perception](<https://devfeed.tech/tags/machine-perception.md>), [models](<https://devfeed.tech/tags/models.md>), [object-detection](<https://devfeed.tech/tags/object-detection.md>), [vision](<https://devfeed.tech/tags/vision.md>)

### AI overview

Google Research presents Google Earth AI, a family of geospatial AI models and reasoning agents that combines foundation models, Gemini-based orchestration, real-world data, datastores, and geospatial tools to answer complex planetary-scale questions. The article also introduces Remote Sensing Foundations models for satellite imagery analysis using vision-language models, open-vocabulary object detection, and adaptable vision backbones.

### Source excerpt

Climate & Sustainability

## VideoPrism: A foundational visual encoder for video understanding

DevFeed: [VideoPrism: A foundational visual encoder for video understanding](<https://devfeed.tech/articles/videoprism-a-foundational-visual-encoder-for-video-understanding-28551.md>)

Original publisher: [Read original article](<http://blog.research.google/2024/02/videoprism-foundational-visual-encoder.html>)

Author: Google AI (noreply@blogger.com)

Published: 2024-02-22T20:05:00Z

Content type: article

Language: en

Sources: [Google Research](<https://devfeed.tech/sources/google-research.md>)

Topics: [foundation-models](<https://devfeed.tech/topics/foundation-models.md>), [dataset](<https://devfeed.tech/topics/dataset.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Google](<https://devfeed.tech/topics/google.md>), [data](<https://devfeed.tech/topics/data.md>), [Software](<https://devfeed.tech/topics/software.md>)

Tags: [classification](<https://devfeed.tech/tags/classification.md>), [dataset](<https://devfeed.tech/tags/dataset.md>), [foundation-models](<https://devfeed.tech/tags/foundation-models.md>), [google](<https://devfeed.tech/tags/google.md>), [machine-intelligence](<https://devfeed.tech/tags/machine-intelligence.md>), [machine-perception](<https://devfeed.tech/tags/machine-perception.md>), [model](<https://devfeed.tech/tags/model.md>), [performance](<https://devfeed.tech/tags/performance.md>), [qa](<https://devfeed.tech/tags/qa.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [training](<https://devfeed.tech/tags/training.md>), [video](<https://devfeed.tech/tags/video.md>)

### AI overview

Google Research introduces VideoPrism, a video foundation model designed for general-purpose video understanding across tasks including classification, localization, retrieval, captioning, and question answering. The model is pretrained on video-text pairs and video clips with noisy or machine-generated text, and the article reports state-of-the-art performance using a single frozen model.

### Source excerpt

Posted by Long Zhao, Senior Research Scientist, and Ting Liu, Senior Staff Software Engineer, Google Research An astounding number of videos are available on the Web, covering a variety of content from everyday moments people share to historical moments to scientific observations, each of which contains a unique record of the world. The right tools could help researchers analyze these videos, transforming how we understand the world around us. Videos offer dynamic visual content far more rich than static images, capturing movement, changes, and dynamic relationships between entities. Analyzing this complexity, along with the immense diversity of publicly available video data, demands models that go beyond traditional image understanding. Consequently, many of the approaches that best perform on video understanding still rely on specialized models tailor-made for particular tasks. Recently, there has been exciting progress in this area using video foundation models (ViFMs), such as VideoCLIP, InternVideo, VideoCoCa, and UMT. However, building a ViFM that handles the sheer diversity of video data remains a challenge. With the goal of building a single model for general-purpose video understanding, we introduce "VideoPrism: A Foundational Visual Encoder for Video Understanding". VideoPrism is a ViFM designed to handle a wide spectrum of video understanding tasks, including classification, localization, retrieval, captioning, and question answering (QA). We propose innovations in both the pre-training data as well as the modeling strategy. We pre-train VideoPrism on a massive and diverse dataset: 36 million high-quality video-text pairs and 582 million video clips with noisy or machine-generated parallel text. Our pre-training approach is designed for this hybrid data, to learn both from video-text pairs and the videos themselves. VideoPrism is incredibly easy to adapt to new video understanding challenges, and achieves state-of-the-art performance using a single frozen