ECCV 2026: Reflections, Highlights, and Key Takeaways

September 15, 2026

TL;DR


Vivek Chavan outside Malmö Arena during ECCV 2026

Back at ECCV, two years after Milan.

I just returned from Malmö after attending ECCV 2026. This was my second ECCV, after previously attending ECCV 2024 in Milan, and the first major research conference that I have now attended twice.

That made this year's event particularly interesting: it was not only a chance to see where the computer vision community is heading, but also to compare how my own experience of participating in these conferences has changed over the last two years.

Vivek Chavan with a workshop poster at ECCV 2026

Five workshop papers presented at ECCV 2026.

I presented our recent and ongoing work across the ECCV workshops. Our contributions covered several closely connected research directions: egocentric and exocentric procedural understanding, neuro-symbolic reasoning, conditional visual grounding for robot learning, industrial egocentric datasets, and multimodal worker assistance.

Two of the papers were selected for oral presentation, at the ACVR and X-Reason workshops.

Workshops: A Snapshot of Emerging Research Directions

The first two days of ECCV were dedicated to workshops and tutorials. The workshop programme was remarkably broad, with more than 90 workshops covering almost every major direction in contemporary computer vision.

The official programme grouped them into themes including 3D Vision & Geometry, Embodied AI, Agents & World Models, Autonomous Driving, Humans, Faces & Behavior, Medical & Biological Vision, Recognition, Segmentation & Video, Generative Models & Content Creation, Multimodal & Foundation Models, Efficiency, Trustworthy & Responsible AI, Earth, Climate & Sustainability, Sensing & Wearables, Art, Culture & Heritage, and Theory & Emerging Directions.

What stood out was how many workshops were no longer centred on a single classical vision task. Instead, many combined perception with larger questions around reasoning, interaction, generation, autonomy, physical understanding, and deployment.

The 3D vision and geometry programme was particularly extensive, covering areas such as open-world 3D scene understanding, SLAM, structure-from-motion, visual localization, digital twins, 3D generation, and geometric intelligence. At the same time, there was a substantial cluster around embodied AI and world models, including interactive agents, dexterous manipulation, multimodal reasoning in physical environments, human-scene interaction, and the evaluation and safety of world models.

Other parts of the programme reflected the increasing breadth of the field. Autonomous-driving workshops focused on foundation models, sim-to-real transfer, and robust autonomy; medical and biological vision covered foundation models, 3D medical imaging, video understanding, and data curation; while workshops on trustworthy AI addressed explainability, fairness, privacy, uncertainty, model unlearning, and visual misinformation.

There was also strong representation from generative and multimodal AI, including audio-visual generation, human-AI co-creation, multimodal language models, evidence-aligned reasoning, and universal representations for perception and world modelling. At the other end of the spectrum, workshops on efficiency, empirical theory, quantum computer vision, ecology, climate, agriculture, cultural heritage, and wearable sensing showed just how widely computer vision is now being applied.

My Focus Areas

Within this much broader programme, my own attention naturally gravitated toward the workshops closest to my research: egocentric and wearable vision, procedural understanding, embodied reasoning, robot learning, and long-horizon vision-language-action systems.

This included workshops such as Assistive Computer Vision and Robotics (ACVR), Visual Perception and Reasoning in the Interactable World (X-Reason), Observing and Acting as Dexterous Hands (DexHAND), Wearables AI, and Foundation Data for Industrial Tech Transfer (FOUND).

Across these sessions, several recurring themes stood out to me. One was the growing connection between vision and action: perception is increasingly being embedded within systems that must reason about task state, make decisions, and interact with the physical world.

Another was the move toward long-horizon understanding. Rather than recognising isolated objects or short actions, many works considered extended activities, temporal structure, memory, procedural state, errors, and recovery. This was especially visible in work around egocentric perception and robotics, where understanding what has already happened can be as important as interpreting the current frame.

A third theme was the growing role of multimodal and first-person sensing. Wearable cameras, language, gaze, audio, and other contextual signals are increasingly being combined to understand human activities and to provide supervision for assistive or robotic systems.

For me, this was one of the main strengths of the workshop programme. The main conference gives a broad picture of where computer vision currently stands; the workshops often provide an earlier view of where specialised research communities are beginning to converge.


The scale of ECCV continues to grow considerably. ECCV 2026 received 10,473 valid submissions from more than 37,000 authors. Of these, 2,834 papers were accepted, corresponding to an acceptance rate of 27.1%. Only 163 papers were selected for oral presentation, or around 1.6% of valid submissions, including 28 long orals and 135 short orals.

The longer-term trend is even more striking. ECCV received 2,439 submissions in 2018, 5,150 in 2020, 5,804 in 2022, 8,585 in 2024, and 10,473 in 2026. In eight years, the number of submissions has grown by more than four times.

My personal impression from the poster halls and oral sessions was that an increasing amount of computer vision research is being connected to larger downstream systems. I repeatedly encountered work on 3D reconstruction, spatial understanding, robotics, embodied AI, autonomous systems, world models, egocentric perception, and multimodal reasoning.

This does not mean that classical computer vision problems are disappearing. Rather, recognition, geometry, tracking, reconstruction, and representation learning are increasingly being placed inside systems that must reason about or interact with the physical world.

The continued strength of 3D vision and geometry was particularly noticeable. This was also reflected in the Best Paper candidate list, which included work on spatial reasoning, streaming 3D reconstruction, geometric representations, structure-from-motion, point-cloud registration, surface normal estimation, and computational imaging.

For a field that is often discussed today mainly through the lens of foundation models and generative AI, ECCV was a useful reminder that understanding the geometry and physical structure of the visual world remains a central research problem.


Keynotes and Invited Talks

ECCV 2026 featured three keynote speakers: Kristen Grauman, Yann LeCun, and Jamie Shotton. The two talks that stood out most to me were those by Kristen Grauman and Yann LeCun.

Kristen Grauman speaking at ECCV 2026
Yann LeCun speaking at ECCV 2026

Kristen Grauman (left) and Yann LeCun (right), two of the keynote speakers at ECCV 2026.

From Machine Perception to Human Intelligence

Speaker: Kristen Grauman

Kristen Grauman’s talk was naturally one of the most relevant to my own research interests. Egocentric vision has developed far beyond the original problem of recognising actions from first-person video. The broader question is increasingly how an intelligent system can understand what a person is doing over long periods of time, relate first-person observations to the wider environment, learn from human demonstrations, and ultimately use that understanding to assist or act.

This direction appeared repeatedly throughout ECCV, not only in Grauman’s keynote but also across workshops and papers on wearable AI, long-video understanding, multimodal perception, and embodied agents.

World Models and a Provocative Research Agenda

Speaker: Yann LeCun

Yann LeCun’s keynote focused on world models and his long-standing argument that today’s dominant AI systems are still missing fundamental capabilities required for human-level intelligence. His emphasis was on systems that learn predictive models of the world and use them for reasoning, planning, and action.

The most memorable moment came near the end, when he gave an unusually direct set of recommendations to his fellow AI scientists. He argued for moving away from generative models in favour of joint-embedding architectures, probabilistic models in favour of energy-based approaches, contrastive methods in favour of regularised methods, and reinforcement learning as the default route to intelligent behaviour in favour of model-predictive control.

His final recommendation was the sharpest of all: researchers interested in human-level AI should not work on LLMs.

Whether one agrees with all of these prescriptions or not, it was refreshing to hear such a clear research position. Much of contemporary AI development is organised around scaling variations of already successful paradigms. LeCun instead argued that several of those paradigms may themselves be dead ends for the longer-term goal of intelligent agents.

That tension between scaling today’s successful systems and searching for fundamentally different architectures was one of the more interesting themes I took away from the conference.

From Vision to Embodied AI

Speaker: Jamie Shotton

The third keynote, by Jamie Shotton, continued another theme visible throughout ECCV: computer vision increasingly serves not only as a mechanism for interpreting images, but as the perceptual foundation of systems that must operate in the physical world.


Best Paper Awards

The ECCV 2026 Best Paper Award was given to:

Heat Kernel Textures – the Geodesic Gaussians That Do Not Splat

Simone Foti, Caner Korkmaz, Stefanos Zafeiriou, Tolga Birdal

The paper introduces an intrinsic representation for texturing triangular meshes using heat kernels on the surface itself. Instead of depending on a conventional UV parameterisation, appearance is represented directly over the mesh geometry.

Two additional papers received Best Paper Honourable Mentions:

Taken together, the recognised papers reinforced something that was already visible throughout the conference: 3D geometry and physical understanding remain major research directions even as the wider AI landscape moves toward increasingly large multimodal models.


Test of Time Awards

ECCV’s Koenderink Prize recognises papers from ten years earlier that have had a lasting impact on computer vision. Among the works recognised from ECCV 2016 were several papers that have become foundational references in modern vision.

Perceptual Losses for Real-Time Style Transfer and Super-Resolution

Justin Johnson, Alexandre Alahi, Li Fei-Fei

This work helped establish the idea of evaluating image similarity using deep feature representations rather than relying only on pixel-level losses. Variants of perceptual losses have since become standard across image synthesis, reconstruction, super-resolution, and generative modelling.

SSD: Single Shot MultiBox Detector

Wei Liu et al.

SSD became one of the defining single-stage detectors of the deep-learning era. By predicting classes and bounding boxes directly from multiple feature-map scales, it helped establish a fast and practical approach to object detection that influenced a decade of subsequent systems.

Learning without Forgetting

Zhizhong Li, Derek Hoiem

This paper is particularly relevant to my own research interests. It addressed catastrophic forgetting: the tendency of a neural network to lose performance on previously learned tasks when adapted to new ones.

A decade later, continual learning remains an active research problem. In fact, the same fundamental question has become even more relevant as increasingly large pretrained models are repeatedly adapted to new domains, skills, and tasks.


Poster Presentations and Oral Sessions

Poster session at ECCV 2026

A lively poster session during the main conference.

As with NeurIPS last year, I found that the poster sessions provided some of the greatest value of the conference.

The sheer number of accepted papers makes it impossible to explore more than a fraction of the programme. Posters solve this surprisingly well. You can move quickly between research areas, stop when something catches your attention, and immediately speak with the people who actually did the work.

There is also a completely different quality to these discussions compared with simply reading the paper. You can ask why an experiment was designed in a certain way, what failed before the final method worked, which limitations the authors are most concerned about, or where they intend to take the work next.

One minor annoyance was the zig-zag arrangement of the poster boards, used during both the workshops and the main conference. Compared with simple straight rows, it made some areas feel unnecessarily cramped and awkward to navigate once groups formed around neighbouring posters. It was a small logistical issue, but a surprisingly noticeable one during the busiest sessions.

Oral presentation session at ECCV 2026

One of the oral presentation sessions at ECCV 2026.

The oral programme offered the opposite experience: a highly compressed selection of work presented to a larger audience. Only 163 papers out of more than ten thousand valid submissions received oral presentations this year, making these sessions a very selective subset of the programme.

The combination worked well. The oral sessions offered a broad snapshot of notable research, while the poster sessions provided the depth and direct interaction. For me, the latter remained more valuable.


Expo and Sponsored Booths

The Expo ran alongside the main conference and brought together major industrial research labs, hardware and software companies, startups, and recruiting teams.

Compared with some academic events, ECCV has an especially visible connection to industrial computer vision research. Many booths went beyond conventional company displays and included live technical demonstrations, research talks, hardware prototypes, and informal opportunities to discuss ongoing work directly with researchers.

Google recruiter giving a talk at the Google booth during ECCV 2026
Meta researcher speaking about DINOv3 at ECCV 2026

A recruiter speaking at the Google booth (left) and a Meta research talk on DINOv3 (right).

I deliberately spent more time exploring the Expo this year. This was something I underestimated at previous conferences. It is easy to focus entirely on the official scientific programme and postpone visiting the booths until later, but the exhibition provides a different cross-section of the field: which problems companies are actively investing in, which technologies are moving from research into products, and which research directions are creating entirely new companies.

The booths were also useful as small presentation spaces in their own right. Alongside recruiting conversations, companies hosted short talks from researchers and technical teams. This made it possible to move between the formal scientific programme and much more informal discussions about current research, engineering, and hiring within the same hall.

The conversations were ultimately at least as valuable as the demonstrations themselves.


General Observations: Venue and Vibe

Malmö made for a very convenient conference location. The conference venues were directly connected to the Hyllie railway station, with Copenhagen and Copenhagen Airport easily reachable across the Öresund connection.

For me, Copenhagen also became part of the conference experience. I spent the week moving between Copenhagen and Malmö, which created the slightly unusual situation of crossing an international border on the way to an academic conference each morning.

Malmö city during ECCV 2026
Railway tracks at sunset in Malmö during ECCV 2026
Snacks and sweets served during an ECCV 2026 poster session

Malmö during ECCV 2026: a scene from the city (left), railway tracks at sunset (centre), and refreshments during a poster session (right).

Inside the conference, the atmosphere was lively throughout the main days. Poster sessions were busy, keynote sessions filled the Arena, and the Expo remained crowded. Despite the size of the event, ECCV still felt more manageable than NeurIPS while being large enough that it was impossible to see everything.


Closing Remarks

ECCV 2026 was particularly meaningful to me because it was the first major research conference I had attended for the second time.

When I attended ECCV in Milan in 2024, I was at a different stage of my research. Returning two years later provided an interesting benchmark. The field has moved, my own research interests have evolved, and I participated in the conference differently this time.

Several areas that are increasingly prominent within computer vision: egocentric perception, long-horizon video understanding, embodied AI, robotics, world models, and reasoning about physical tasks, now overlap strongly with the questions I work on myself.

At the same time, the biggest takeaway from actually travelling to these conferences remains remarkably consistent.

Most papers can be read online. Slides can be downloaded. Talks increasingly appear as recordings. What remains difficult to reproduce remotely are the spontaneous conversations: finding a poster you had not planned to visit, asking an author one very specific question, meeting somebody working on the same problem from a completely different direction, or continuing a technical discussion long after the formal session has ended.

Those interactions were the most valuable part of ECCV for me.

Two years after Milan, it was exciting to return and see both how much the community has changed and how much more there is still to explore.

← Back to Blog


💬 Comments

Comments are powered by Utterances. You’ll need a GitHub account to post.