Ross Girshick
Meta (United States)
14 papers in the Anchorcite directory, published between 2009 and 2023, cited 334,262 times in total.
Research topics
- Advanced Neural Network Applications · 14 papers
- Advanced Image and Video Retrieval Techniques · 8 papers
- Domain Adaptation and Few-Shot Learning · 6 papers
- Video Surveillance and Tracking Methods · 4 papers
- Visual Attention and Saliency Detection · 2 papers
- Human Pose and Action Recognition · 2 papers
Frequent co-authors
- Kaiming He · 8 papers
- Piotr Dollár · 6 papers
- Tsung-Yi Lin · 3 papers
- Jeff Donahue · 2 papers
- Trevor J. Darrell · 2 papers
- Priya Goyal · 2 papers
- Saining Xie · 2 papers
- Shaoqing Ren · 1 paper
Papers
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal NetworksIEEE Transactions on Pattern Analysis and Machine Intelligence · 2016 · 56,050 citations
State-of-the-art object detection networks depend on region proposal algorithms to hypothesize object locations. Advances like SPPnet [1] and Fast R-CNN [2] have reduced the running time of these detection networks, exposing region proposal computation as a bottleneck. In this work,…
- You Only Look Once: Unified, Real-Time Object Detection2016 · 51,627 citations
We present YOLO, a new approach to object detection. Prior work on object detection repurposes classifiers to perform detection. Instead, we frame object detection as a regression problem to spatially separated bounding boxes and associated class probabilities. A single neural…
- Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation2014 · 32,278 citations
Object detection performance, as measured on the canonical PASCAL VOC dataset, has plateaued in the last few years. The best-performing methods are complex ensemble systems that typically combine multiple low-level image features with high-level context. In this paper, we propose…
- Feature Pyramid Networks for Object Detection2017 · 30,041 citations
Feature pyramids are a basic component in recognition systems for detecting objects at different scales. But pyramid representations have been avoided in recent object detectors that are based on deep convolutional networks, partially because they are slow to compute and…
- Mask R-CNN2017 · 29,845 citations
We present a conceptually simple, flexible, and general framework for object instance segmentation. Our approach efficiently detects objects in an image while simultaneously generating a high-quality segmentation mask for each instance. The method, called Mask R-CNN, extends Faster R-CNN by…
- Fast R-CNN2015 · 28,653 citations
This paper proposes a Fast Region-based Convolutional Network method (Fast R-CNN) for object detection. Fast R-CNN builds on previous work to efficiently classify object proposals using deep convolutional networks. Compared to previous work, Fast R-CNN employs several innovations to improve…
- Focal Loss for Dense Object Detection2017 · 27,631 citations
The highest accuracy object detectors to date are based on a two-stage approach popularized by R-CNN, where a classifier is applied to a sparse set of candidate object locations. In contrast, one-stage detectors that are applied over a regular, dense…
- Momentum Contrast for Unsupervised Visual Representation Learning2020 · 12,628 citations
We present Momentum Contrast (MoCo) for unsupervised visual representation learning. From a perspective on contrastive learning as dictionary look-up, we build a dynamic dictionary with a queue and a moving-averaged encoder. This enables building a large and consistent dictionary on-the-fly…
- Aggregated Residual Transformations for Deep Neural Networks2017 · 12,052 citations
We present a simple, highly modularized network architecture for image classification. Our network is constructed by repeating a building block that aggregates a set of transformations with the same topology. Our simple design results in a homogeneous, multi-branch architecture that…
- Non-local Neural Networks2018 · 11,441 citations
Both convolutional and recurrent operations are building blocks that process one local neighborhood at a time. In this paper, we present non-local operations as a generic family of building blocks for capturing long-range dependencies. Inspired by the classical non-local means…
- Caffe2014 · 11,223 citations
Caffe provides multimedia scientists and practitioners with a clean and modifiable framework for state-of-the-art deep learning algorithms and a collection of reference models. The framework is a BSD-licensed C++ library with Python and MATLAB bindings for training and deploying general-purpose…
- Segment Anything2023 · 10,762 citations
We introduce the Segment Anything (SA) project: a new task, model, and dataset for image segmentation. Using our efficient model in a data collection loop, we built the largest segmentation dataset to date (by far), with over 1 billion masks…
- Object Detection with Discriminatively Trained Part-Based ModelsIEEE Transactions on Pattern Analysis and Machine Intelligence · 2009 · 10,131 citations
We describe an object detection system based on mixtures of multiscale deformable part models. Our system is able to represent highly variable object classes and achieves state-of-the-art results in the PASCAL object detection challenges. While deformable part models have become…
- Focal Loss for Dense Object DetectionIEEE Transactions on Pattern Analysis and Machine Intelligence · 2018 · 9,900 citations
The highest accuracy object detectors to date are based on a two-stage approach popularized by R-CNN, where a classifier is applied to a sparse set of candidate object locations. In contrast, one-stage detectors that are applied over a regular, dense…
Metadata from OpenAlex (CC0).