Kaiming He
Meta (United States)
11 papers in the Anchorcite directory, published between 2015 and 2020, cited 449,679 times in total.
ORCID profileResearch topics
- Advanced Neural Network Applications · 11 papers
- Domain Adaptation and Few-Shot Learning · 8 papers
- Advanced Image and Video Retrieval Techniques · 5 papers
- Video Surveillance and Tracking Methods · 3 papers
- Human Pose and Action Recognition · 2 papers
- Adversarial Robustness in Machine Learning · 2 papers
Frequent co-authors
- Ross Girshick · 8 papers
- Piotr Dollár · 5 papers
- Shaoqing Ren · 4 papers
- Xiangyu Zhang · 3 papers
- Tsung-Yi Lin · 3 papers
- Jian Sun · 2 papers
- Priya Goyal · 2 papers
- Saining Xie · 2 papers
Papers
- Deep Residual Learning for Image Recognition2016 · 229,413 citations
Deeper neural networks are more difficult to train. We present a residual learning framework to ease the training of networks that are substantially deeper than those used previously. We explicitly reformulate the layers as learning residual functions with reference to…
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal NetworksIEEE Transactions on Pattern Analysis and Machine Intelligence · 2016 · 56,050 citations
State-of-the-art object detection networks depend on region proposal algorithms to hypothesize object locations. Advances like SPPnet [1] and Fast R-CNN [2] have reduced the running time of these detection networks, exposing region proposal computation as a bottleneck. In this work,…
- Feature Pyramid Networks for Object Detection2017 · 30,041 citations
Feature pyramids are a basic component in recognition systems for detecting objects at different scales. But pyramid representations have been avoided in recent object detectors that are based on deep convolutional networks, partially because they are slow to compute and…
- Mask R-CNN2017 · 29,845 citations
We present a conceptually simple, flexible, and general framework for object instance segmentation. Our approach efficiently detects objects in an image while simultaneously generating a high-quality segmentation mask for each instance. The method, called Mask R-CNN, extends Faster R-CNN by…
- Focal Loss for Dense Object Detection2017 · 27,631 citations
The highest accuracy object detectors to date are based on a two-stage approach popularized by R-CNN, where a classifier is applied to a sparse set of candidate object locations. In contrast, one-stage detectors that are applied over a regular, dense…
- Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification2015 · 19,060 citations
Rectified activation units (rectifiers) are essential for state-of-the-art neural networks. In this work, we study rectifier neural networks for image classification from two aspects. First, we propose a Parametric Rectified Linear Unit (PReLU) that generalizes the traditional rectified unit. PReLU…
- Momentum Contrast for Unsupervised Visual Representation Learning2020 · 12,628 citations
We present Momentum Contrast (MoCo) for unsupervised visual representation learning. From a perspective on contrastive learning as dictionary look-up, we build a dynamic dictionary with a queue and a moving-averaged encoder. This enables building a large and consistent dictionary on-the-fly…
- Aggregated Residual Transformations for Deep Neural Networks2017 · 12,052 citations
We present a simple, highly modularized network architecture for image classification. Our network is constructed by repeating a building block that aggregates a set of transformations with the same topology. Our simple design results in a homogeneous, multi-branch architecture that…
- Spatial Pyramid Pooling in Deep Convolutional Networks for Visual RecognitionIEEE Transactions on Pattern Analysis and Machine Intelligence · 2015 · 11,618 citations
Existing deep convolutional neural networks (CNNs) require a fixed-size (e.g., 224 × 224) input image. This requirement is "artificial" and may reduce the recognition accuracy for the images or sub-images of an arbitrary size/scale. In this work, we equip the…
- Non-local Neural Networks2018 · 11,441 citations
Both convolutional and recurrent operations are building blocks that process one local neighborhood at a time. In this paper, we present non-local operations as a generic family of building blocks for capturing long-range dependencies. Inspired by the classical non-local means…
- Focal Loss for Dense Object DetectionIEEE Transactions on Pattern Analysis and Machine Intelligence · 2018 · 9,900 citations
The highest accuracy object detectors to date are based on a two-stage approach popularized by R-CNN, where a classifier is applied to a sparse set of candidate object locations. In contrast, one-stage detectors that are applied over a regular, dense…
Metadata from OpenAlex (CC0).