SingDancer

常用图像数据集大全（分类，跟踪，分割，检测等）

1.搜狗实验室数据集：

http://www.sogou.com/labs/dl/p.html

互联网图片库来自sogou图片搜索所索引的部分数据。其中收集了包括人物、动物、建筑、机械、风景、运动等类别，总数高达2,836,535张图片。对于每张图片，数据集中给出了图片的原图、缩略图、所在网页以及所在网页中的相关文本。200多G

http://www.imageclef.org/

IMAGECLEF致力于位图片相关领域提供一个基准（检索、分类、标注等等） Cross Language Evaluation Forum (CLEF) 。从2003年开始每年举行一次比赛.

http://staff.science.uva.nl/~xirong/index.php?n=Main.Dataset

Xiaorong Li 维护的数据集。PhD ,Intelligent Systems Lab Amsterdam.research on video and image retrieval.

Flickr-3.5M: A collection of 3.5 million social-tagged images.
Social20: A ground-truth set for tag-based social image retrieval.
Biconcepts2012test: A ground-truth set for retrieving bi-concepts (concept pairs) in unlabeled images.
neg4free: A set of negative examples automatically harvested from social-tagged images for 20 PASCAL VOC concepts.

wikipedia featured articles 函数图片（以及特征）以及对应的wiki文本。可以看看文章A New Approach to Cross-Modal Multimedia Retrieval，还有一批文章On the Role of Correlation and Abstraction in Cross-Modal Multimedia Retrieval不过还没有下载链接

http://www.svcl.ucsd.edu/projects/crossmodal/

http://lms.comp.nus.edu.sg/research/NUS-WIDE.htm

To our knowledge, this is the largest real-world web image dataset comprising over 269,000 images with over 5,000 user-provided tags, and ground-truth of 81 concepts for the entire dataset. The dataset is much larger than the popularly available Corel and Caltech 101 datasets. Though some datasets comprise over 3 million images, they only have ground-truth for a small fraction of images. Our proposed NUS-WIDE dataset has the ground-truth for the entire dataset.

http://www.cs.washington.edu/research/imagedatabase/

http://lear.inrialpes.fr/~jegou/data.php

Jegou的数据集，不过Jegou是专门做CBIR的，图像有ground truth，没有标注。

http://www.robots.ox.ac.uk/~vgg/data/oxbuildings/

vgg的osford building dataset。也是专门CBIR的数据。

http://acmmm13.org/submissions/call-for-multimedia-grand-challenge-solutions/msr-bing-grand-challenge-on-image-retrieval-scientific-track/

The dataset for the Microsoft Image Grand Challenge on Image Retrieval

另外介绍cvpaper上的整理的数据集

http://www.cvpapers.com/index.html

Participate in Reproducible Research

Detection

PASCAL VOC 2009 dataset: Classification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets
LabelMe dataset: LabelMe is a web-based image annotation tool that allows researchers to label images and share the annotations with the rest of the community. If you use the database, we only ask that you contribute to it, from time to time, by using the labeling tool.
BioID Face Detection Database: 1521 images with human faces, recorded under natural conditions, i.e. varying illumination and complex background. The eye positions have been set manually.
CMU/VASC & PIE Face dataset
Yale Face dataset
Caltech: Cars, Motorcycles, Airplanes, Faces, Leaves, Backgrounds
Caltech 101: Pictures of objects belonging to 101 categories
Caltech 256: Pictures of objects belonging to 256 categories
Daimler Pedestrian Detection Benchmark: 15,560 pedestrian and non-pedestrian samples (image cut-outs) and 6744 additional full images not containing pedestrians for bootstrapping. The test set contains more than 21,790 images with 56,492 pedestrian labels (fully visible or partially occluded), captured from a vehicle in urban traffic.
MIT Pedestrian dataset: CVC Pedestrian Datasets
CVC Pedestrian Datasets: CBCL Pedestrian Database
MIT Face dataset: CBCL Face Database
MIT Car dataset: CBCL Car Database
MIT Street dataset: CBCL Street Database
INRIA Person Data Set: A large set of marked up images of standing or walking people
INRIA car dataset: A set of car and non-car images taken in a parking lot nearby INRIA
INRIA horse dataset: A set of horse and non-horse images
H3D Dataset: 3D skeletons and segmented regions for 1000 people in images
HRI RoadTraffic dataset: A large-scale vehicle detection dataset
BelgaLogos: 10000 images of natural scenes, with 37 different logos, and 2695 logos instances, annotated with a bounding box.
FlickrBelgaLogos: 10000 images of natural scenes grabbed on Flickr, with 2695 logos instances cut and pasted from the BelgaLogos dataset.
FlickrLogos-32: The dataset FlickrLogos-32 contains photos depicting logos and is meant for the evaluation of multi-class logo detection/recognition as well as logo retrieval methods on real-world images. It consists of 8240 images downloaded from Flickr.
TME Motorway Dataset: 30000+ frames with vehicle rear annotation and classification (car and trucks) on motorway/highway sequences. Annotation semi-automatically generated using laser-scanner data. Distance estimation and consistent target ID over time available.
PHOS (Color Image Database for illumination invariant feature selection): Phos is a color image database of 15 scenes captured under different illumination conditions. More particularly, every scene of the database contains 15 different images: 9 images captured under various strengths of uniform illumination, and 6 images under different degrees of non-uniform illumination. The images contain objects of different shape, color and texture and can be used for illumination invariant feature detection and selection.
CaliforniaND: An Annotated Dataset For Near-Duplicate Detection In Personal Photo Collections: California-ND contains 701 photos taken directly from a real user's personal photo collection, including many challenging non-identical near-duplicate cases, without the use of artificial image transformations. The dataset is annotated by 10 different subjects, including the photographer, regarding near duplicates.

Classification

PASCAL VOC 2009 dataset: Classification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets
Caltech: Cars, Motorcycles, Airplanes, Faces, Leaves, Backgrounds
Caltech 101: Pictures of objects belonging to 101 categories
Caltech 256: Pictures of objects belonging to 256 categories
ETHZ Shape Classes: A dataset for testing object class detection algorithms. It contains 255 test images and features five diverse shape-based classes (apple logos, bottles, giraffes, mugs, and swans).
Flower classification data sets: 17 Flower Category Dataset
Animals with attributes: A dataset for Attribute Based Classification. It consists of 30475 images of 50 animals classes with six pre-extracted feature representations for each image.
Stanford Dogs Dataset: Dataset of 20,580 images of 120 dog breeds with bounding-box annotation, for fine-grained image categorization.

Recognition

Face and Gesture Recognition Working Group FGnet: Face and Gesture Recognition Working Group FGnet
Feret: Face and Gesture Recognition Working Group FGnet
PUT face: 9971 images of 100 people
Labeled Faces in the Wild: A database of face photographs designed for studying the problem of unconstrained face recognition
Urban scene recognition: Traffic Lights Recognition, Lara's public benchmarks.
PubFig: Public Figures Face Database: The PubFig database is a large, real-world face dataset consisting of 58,797 images of 200 people collected from the internet. Unlike most other existing face datasets, these images are taken in completely uncontrolled situations with non-cooperative subjects.
YouTube Faces: The data set contains 3,425 videos of 1,595 different people. The shortest clip duration is 48 frames, the longest clip is 6,070 frames, and the average length of a video clip is 181.3 frames.
MSRC-12: Kinect gesture data set: The Microsoft Research Cambridge-12 Kinect gesture data set consists of sequences of human movements, represented as body-part locations, and the associated gesture to be recognized by the system.
QMUL underGround Re-IDentification (GRID) Dataset: This dataset contains 250 pedestrian image pairs + 775 additional images captured in a busy underground station for the research on person re-identification.
Person identification in TV series: Face tracks, features and shot boundaries from our latest CVPR 2013 paper. It is obtained from 6 episodes of Buffy the Vampire Slayer and 6 episodes of Big Bang Theory.
ChokePoint Dataset: ChokePoint is a video dataset designed for experiments in person identification/verification under real-world surveillance conditions. The dataset consists of 25 subjects (19 male and 6 female) in portal 1 and 29 subjects (23 male and 6 female) in portal 2.

Tracking

BIWI Walking Pedestrians dataset: Walking pedestrians in busy scenarios from a bird eye view
"Central" Pedestrian Crossing Sequences: Three pedestrian crossing sequences
Pedestrian Mobile Scene Analysis: The set was recorded in Zurich, using a pair of cameras mounted on a mobile platform. It contains 12'298 annotated pedestrians in roughly 2'000 frames.
Head tracking: BMP image sequences.
KIT AIS Dataset: Data sets for tracking vehicles and people in aerial image sequences.
MIT Traffic Data Set: MIT traffic data set is for research on activity analysis and crowded scenes. It includes a traffic video sequence of 90 minutes long. It is recorded by a stationary camera.

Segmentation

Image Segmentation with A Bounding Box Prior dataset: Ground truth database of 50 images with: Data, Segmentation, Labelling - Lasso, Labelling - Rectangle
PASCAL VOC 2009 dataset: Classification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets
Motion Segmentation and OBJCUT data: Cows for object segmentation, Five video sequences for motion segmentation
Geometric Context Dataset: Geometric Context Dataset: pixel labels for seven geometric classes for 300 images
Crowd Segmentation Dataset: This dataset contains videos of crowds and other high density moving objects. The videos are collected mainly from the BBC Motion Gallery and Getty Images website. The videos are shared only for the research purposes. Please consult the terms and conditions of use of these videos from the respective websites.
CMU-Cornell iCoseg Dataset: Contains hand-labelled pixel annotations for 38 groups of images, each group containing a common foreground. Approximately 17 images per group, 643 images total.
Segmentation evaluation database: 200 gray level images along with ground truth segmentations
The Berkeley Segmentation Dataset and Benchmark: Image segmentation and boundary detection. Grayscale and color segmentations for 300 images, the images are divided into a training set of 200 images, and a test set of 100 images.
Weizmann horses: 328 side-view color images of horses that were manually segmented. The images were randomly collected from the WWW.
Saliency-based video segmentation with sequentially updated priors: 10 videos as inputs, and segmented image sequences as ground-truth

Foreground/Background

Wallflower Dataset: For evaluating background modelling algorithms
Foreground/Background Microsoft Cambridge Dataset: Foreground/Background segmentation and Stereo dataset from Microsoft Cambridge
Stuttgart Artificial Background Subtraction Dataset: The SABS (Stuttgart Artificial Background Subtraction) dataset is an artificial dataset for pixel-wise evaluation of background models.

Saliency Detection (source)

AIM: 120 Images / 20 Observers (Neil D. B. Bruce and John K. Tsotsos 2005).
LeMeur: 27 Images / 40 Observers (O. Le Meur, P. Le Callet, D. Barba and D. Thoreau 2006).
Kootstra: 100 Images / 31 Observers (Kootstra, G., Nederveen, A. and de Boer, B. 2008).
DOVES: 101 Images / 29 Observers (van der Linde, I., Rajashekar, U., Bovik, A.C., Cormack, L.K. 2009).
Ehinger: 912 Images / 14 Observers (Krista A. Ehinger, Barbara Hidalgo-Sotelo, Antonio Torralba and Aude Oliva 2009).
NUSEF: 758 Images / 75 Observers (R. Subramanian, H. Katti, N. Sebe1, M. Kankanhalli and T-S. Chua 2010).
JianLi: 235 Images / 19 Observers (Jian Li, Martin D. Levine, Xiangjing An and Hangen He 2011).
Extended Complex Scene Saliency Dataset (ECSSD): ECSSD contains 1000 natural images with complex foreground or background. For each image, the ground truth mask of salient object(s) is provided.

Video Surveillance

CAVIAR: For the CAVIAR project a number of video clips were recorded acting out the different scenarios of interest. These include people walking alone, meeting with others, window shopping, entering and exitting shops, fighting and passing out and last, but not least, leaving a package in a public place.
ViSOR: ViSOR contains a large set of multimedia data and the corresponding annotations.

Multiview

3D Photography Dataset: Multiview stereo data sets: a set of images
Multi-view Visual Geometry group's data set: Dinosaur, Model House, Corridor, Aerial views, Valbonne Church, Raglan Castle, Kapel sequence
Oxford reconstruction data set (building reconstruction): Oxford colleges
Multi-View Stereo dataset (Vision Middlebury): Temple, Dino
Multi-View Stereo for Community Photo Collections: Venus de Milo, Duomo in Pisa, Notre Dame de Paris
IS-3D Data: Dataset provided by Center for Machine Perception
CVLab dataset: CVLab dense multi-view stereo image database
3D Objects on Turntable: Objects viewed from 144 calibrated viewpoints under 3 different lighting conditions
Object Recognition in Probabilistic 3D Scenes: Images from 19 sites collected from a helicopter flying around Providence, RI. USA. The imagery contains approximately a full circle around each site.
Multiple cameras fall dataset: 24 scenarios recorded with 8 IP video cameras. The first 22 first scenarios contain a fall and confounding events, the last 2 ones contain only confounding events.

Action

UCF Sports Action Dataset: This dataset consists of a set of actions collected from various sports which are typically featured on broadcast television channels such as the BBC and ESPN. The video sequences were obtained from a wide range of stock footage websites including BBC Motion gallery, and GettyImages.
UCF Aerial Action Dataset: This dataset features video sequences that were obtained using a R/C-controlled blimp equipped with an HD camera mounted on a gimbal.The collection represents a diverse pool of actions featured at different heights and aerial viewpoints. Multiple instances of each action were recorded at different flying altitudes which ranged from 400-450 feet and were performed by different actors.
UCF YouTube Action Dataset: It contains 11 action categories collected from YouTube.
Weizmann action recognition: Walk, Run, Jump, Gallop sideways, Bend, One-hand wave, Two-hands wave, Jump in place, Jumping Jack, Skip.
UCF50: UCF50 is an action recognition dataset with 50 action categories, consisting of realistic videos taken from YouTube.
ASLAN: The Action Similarity Labeling (ASLAN) Challenge.
MSR Action Recognition Datasets: The dataset was captured by a Kinect device. There are 12 dynamic American Sign Language (ASL) gestures, and 10 people. Each person performs each gesture 2-3 times.
KTH Recognition of human actions: Contains six types of human actions (walking, jogging, running, boxing, hand waving and hand clapping) performed several times by 25 subjects in four different scenarios: outdoors, outdoors with scale variation, outdoors with different clothes and indoors.
Hollywood-2 Human Actions and Scenes dataset: Hollywood-2 datset contains 12 classes of human actions and 10 classes of scenes distributed over 3669 video clips and approximately 20.1 hours of video in total.
Collective Activity Dataset: This dataset contains 5 different collective activities : crossing, walking, waiting, talking, and queueing and 44 short video sequences some of which were recorded by consumer hand-held digital camera with varying view point.
Olympic Sports Dataset: The Olympic Sports Dataset contains YouTube videos of athletes practicing different sports.
SDHA 2010: Surveillance-type videos
VIRAT Video Dataset: The dataset is designed to be realistic, natural and challenging for video surveillance domains in terms of its resolution, background clutter, diversity in scenes, and human activity/event categories than existing action recognition datasets.
HMDB: A Large Video Database for Human Motion Recognition: Collected from various sources, mostly from movies, and a small proportion from public databases, YouTube and Google videos. The dataset contains 6849 clips divided into 51 action categories, each containing a minimum of 101 clips.
Stanford 40 Actions Dataset: Dataset of 9,532 images of humans performing 40 different actions, annotated with bounding-boxes.
50Salads dataset: Fully annotated dataset of RGB-D video data and data from accelerometers attached to kitchen objects capturing 25 people preparing two mixed salads each (4.5h of annotated data). Annotated activities correspond to steps in the recipe and include phase (pre-/ core-/ post) and the ingredient acted upon.

Human pose/Expression

AFEW (Acted Facial Expressions In The Wild)/SFEW (Static Facial Expressions In The Wild): Dynamic temporal facial expressions data corpus consisting of close to real world environment extracted from movies.
ETHZ CALVIN Dataset

Image stitching

IPM Vision Group Image Stitching datasets: Images and parameters for registeration

Medical

VIP Laparoscopic / Endoscopic Dataset: Collection of endoscopic and laparoscopic (mono/stereo) videos and images

Misc

Zurich Buildings Database: ZuBuD Image Database contains over 1005 images about Zurich city building.
Color Name Data Sets
Mall dataset: The mall dataset was collected from a publicly accessible webcam for crowd counting and activity profiling research.
QMUL Junction Dataset: A busy traffic dataset for research on activity analysis and behaviour understanding.

CVOnline的数据集

http://homepages.inf.ed.ac.uk/rbf/CVonline/CVentry.htm

Index by Topic

Action Databases
Biological/Medical
Face Databases
Fingerprints
General Images
Gesture Databases
Image, Video and Shape Database Retrieval
Object Databases
People, Pedestrian, Eye/Iris, Template Detection/Tracking Databases
Segmentation
Surveillance
Textures
General Videos
Other Collection Pages
Miscellaneous Topics

Action Databases

50 Salads - fully annotated 4.5 hour dataset of RGB-D video + accelerometer data, capturing 25 people preparing two mixed salads each (Dundee University, Sebastian Stein)
ASLAN Action similarity labeling challenge database (Orit Kliper-Gross)
Berkeley MHAD: A Comprehensive Multimodal Human Action Database (Ferda Ofli)
BEHAVE Interacting Person Video Data with markup (Scott Blunsden, Bob Fisher, Aroosha Laghaee)
CVBASE06: annotated sports videos (Janez Pers)
G3D - synchronised video, depth and skeleton data for 20 gaming actions captured with Microsoft Kinect (Victoria Bloom)
Hollywood 3D - 650 3D action recognition in the wild videos, 14 action classes (Simon Hadfield)
Human Actions and Scenes Dataset (Marcin Marszalek, Ivan Laptev, Cordelia Schmid)
HumanEva: Synchronized Video and Motion Capture Dataset for Evaluation of Articulated Human Motion (Brown University)
i3DPost Multi-View Human Action Datasets (Hansung Kim)
i-LIDS video event image dataset (Imagery library for intelligent detection systems) (Paul Hosner)
INRIA Xmas Motion Acquisition Sequences (IXMAS) (INRIA)
JPL First-Person Interaction dataset - 7 types of human activity videos taken from a first-person viewpoint (Michael S. Ryoo, JPL)
KTH human action recognition database (KTH CVAP lab)
LIRIS human activities dataset - 2 cameras, annotated, depth images (Christian Wolf, et al)
MuHAVi - Multicamera Human Action Video Data (Hossein Ragheb)
Oxford TV based human interactions (Oxford Visual Geometry Group)
Rochester Activities of Daily Living Dataset (Ross Messing)
SDHA Semantic Description of Human Activities 2010 contest - aerial views (Michael S. Ryoo, J. K. Aggarwal, Amit K. Roy-Chowdhury)
SDHA Semantic Description of Human Activities 2010 contest - Human Interactions (Michael S. Ryoo, J. K. Aggarwal, Amit K. Roy-Chowdhury)
TUM Kitchen Data Set of Everyday Manipulation Activities (Moritz Tenorth, Jan Bandouch)
TV Human Interaction Dataset (Alonso Patron-Perez)
Univ of Central Florida - Feature Films Action Dataset (Univ of Central Florida)
Univ of Central Florida - YouTube Action Dataset (sports) (Univ of Central Florida)
Univ of Central Florida - 50 Action Category Recognition in Realistic Videos (3 GB) (Kishore Reddy)
UCF 101 action dataset 101 action classes, over 13k clips and 27 hours of video data (Univ of Central Florida)
Univ of Central Florida - Sports Action Dataset (Univ of Central Florida)
Univ of Central Florida - ARG Aerial camera, Rooftop camera and Ground camera (UCF Computer Vision Lab)
UCR Videoweb Multi-camera Wide-Area Activities Dataset (Amit K. Roy-Chowdhury)
Verona Social interaction dataset (Marco Cristani)
Videoweb (multicamera) Activities Dataset (B. Bhanu, G. Denina, C. Ding, A. Ivers, A. Kamal, C. Ravishankar, A. Roy-Chowdhury, B. Varda)
ViHASi: Virtual Human Action Silhouette Data (userID: VIHASI password: virtual$virtual) (Hossein Ragheb, Kingston University)
WorkoutSU-10 Kinect dataset for exercise actions (Ceyhun Akgul)
YouCook - 88 open-source YouTube cooking videos with annotations (Jason Corso)
WVU Multi-view action recognition dataset (Univ. of West Virginia)

Biological/Medical

Computed Tomography Emphysema Database (Lauge Sorensen)
Dermoscopy images (Eric Ehrsam)
DIADEM: Digital Reconstruction of Axonal and Dendritic Morphology Competition (Allen Institute for Brain Science et al)
DIARETDB1 - Standard Diabetic Retinopathy Database (Lappeenranta Univ of Technology)
DRIVE: Digital Retinal Images for Vessel Extraction (Univ of Utrecht)
MiniMammographic Database (Mammographic Image Analysis Society)
MIT CBCL Automated Mouse Behavior Recognition datasets (Nicholas Edelman)
Retinal fundus images - Ground truth of vascular bifurcations and crossovers (Univ of Groningen)
Spine and Cardiac data (Digital Imaging Group of London Ontario, Shuo Li)
Univ of Central Florida - DDSM: Digital Database for Screening Mammography (Univ of Central Florida)
VascuSynth - 120 3D vascular tree like structures with ground truth (Mengliu Zhao, Ghassan Hamarneh)
York Cardiac MRI dataset (Alexander Andreopoulos)

Face Databases

3D Mask Attack Database (3DMAD) - 76500 frames of 17 persons using Kinect RGBD with eye positions (Sebastien Marcel)
Audio-visual database for face and speaker recognition (Mobile Biometry MOBIO http://www.mobioproject.org/)
BANCA face and voice database (Univ of Surrey)
Binghampton Univ 3D static and dynamic facial expression database (Lijun Yin, Peter Gerhardstein and teammates)
BioID face database (BioID group)
Biwi 3D Audiovisual Corpus of Affective Communication - 1000 high quality, dynamic 3D scans of faces, recorded while pronouncing a set of English sentences.
CMU Facial Expression Database (CMU/MIT)
CMU/MIT Frontal Faces (CMU/MIT)
CMU/MIT Frontal Faces (CMU/MIT)
CMU Pose, Illumination, and Expression (PIE) Database (Simon Baker)
CSSE Frontal intensity and range images of faces (Ajmal Mian)
Face Recognition Grand Challenge datasets (FRVT - Face Recognition Vendor Test)
FaceTracer Database - 15,000 faces (Neeraj Kumar, P. N. Belhumeur, and S. K. Nayar)
FDDB: Face Detection Data set and Benchmark - studying unconstrained face detection (University of Massachusetts Computer Vision Laboratory)
FG-Net Aging Database of faces at different ages (Face and Gesture Recognition Research Network)
Facial Recognition Technology (FERET) Database (USA National Institute of Standards and Technology)
Hong Kong Face Sketch Database
Japanese Female Facial Expression (JAFFE) Database (Michael J. Lyons)
LFW: Labeled Faces in the Wild - unconstrained face recognition. Re-labeled Faces in the Wild - original images, but aligned using "deep funneling" method. (University of Massachusetts, Amherst)
Manchester Annotated Talking Face Video Dataset (Timothy Cootes)
MIT Collation of Face Databases (Ethan Meyers)
MORPH (Craniofacial Longitudinal Morphological Face Database) (University of North Carolina Wilmington)
MIT CBCL Face Recognition Database (Center for Biological and Computational Learning)
NIST mugshot identification database (USA National Institute of Standards and Technology)
ORL face database: 40 people with 10 views (ATT Cambridge Labs)
Oxford: faces, flowers, multi-view, buildings, object categories, motion segmentation, affine covariant regions, misc (Oxford Visual Geometry Group)
PubFig: Public Figures Face Database (Neeraj Kumar, Alexander C. Berg, Peter N. Belhumeur, and Shree K. Nayar)
SCface - Surveillance Cameras Face Database (Mislav Grgic, Kresimir Delac, Sonja Grgic, Bozidar Klimpak))
Trondheim Kinect RGB-D Person Re-identification Dataset (Igor Barros Barbosa)
UB KinFace Database - University of Buffalo kinship verification and recognition database
XM2VTS Face video sequences (295): The extended M2VTS Database (XM2VTS) - (Surrey University)
Yale Face Database - 11 expressions of 10 people (A. Georghaides)
Yale Face Database B - 576 viewing conditions of 10 people (A. Georghaides)

Fingerprints

FVC fingerpring verification competition 2002 dataset (University of Bologna)
FVC fingerpring verification competition 2004 dataset (University of Bologna)
FVC - a subset of FVC (Fingerprint Verification Competition) 2002 and 2004 fingerprint image databases, manually extracted minutiae data & associated documents (Umut Uludag)
NIST fingerprint databases (USA National Institute of Standards and Technology)
SPD2010 Fingerprint Singular Points Detection Competition (SPD 2010 committee)

General Images

Aerial color image dataset (Swiss Federal Institute of Technology)
AMOS: Archive of Many Outdoor Scenes (20+m) (Nathan Jacobs)
Brown Univ Large Binary Image Database (Ben Kimia)
Columbia Multispectral Image Database (F. Yasuma, T. Mitsunaga, D. Iso, and S.K. Nayar)
HIPR2 Image Catalogue of different types of images (Bob Fisher et al)
Hyperspectral images of natural scenes - 2002 (David H. Foster)
Hyperspectral images of natural scenes - 2004 (David H. Foster)
ImageNet Linguistically organised (WordNet) Hierarchical Image Database - 10E7 images, 15K categories (Li Fei-Fei, Jia Deng, Hao Su, Kai Li)
ImageNet Large Scale Visual Recognition Challenge (Alex Berg, Jia Deng, Fei-Fei Li)
OTCBVS Thermal Imagery Benchmark Dataset Collection (Ohio State Team)
McGill Calibrated Colour Image Database (Adriana Olmos and Fred Kingdom)
Tiny Images Dataset 79 million 32x32 color images (Fergus, Torralba, Freeman)

Gesture Databases

FG-Net Aging Database of faces at different ages (Face and Gesture Recognition Research Network)
Hand gesture and marine silhouettes (Euripides G.M. Petrakis)
IDIAP Hand pose/gesture datasets (Sebastien Marcel)
Sheffield gesture database - 2160 RGBD hand gesture sequences, 6 subjects, 10 gestures, 3 postures, 3 backgrounds, 2 illuminations (Ling Shao)

Image, Video and Shape Database Retrieval

Brown Univ 25/99/216 Shape Databases (Ben Kimia)
IAPR TC-12 Image Benchmark (Michael Grubinger)
IAPR-TC12 Segmented and annotated image benchmark (SAIAPR TC-12): (Hugo Jair Escalante)
ImageCLEF 2010 Concept Detection and Annotation Task (Stefanie Nowak)
ImageCLEF 2011 Concept Detection and Annotation Task - multi-label classification challenge in Flickr photos
CLEF-IP 2011 evaluation on patent images
McGill 3D Shape Benchmark (Siddiqi, Zhang, Macrini, Shokoufandeh, Bouix, Dickinson)
NIST SHREC 2010 - Shape Retrieval Contest of Non-rigid 3D Models (USA National Institute of Standards and Technology)
NIST SHREC - other NIST retrieval contest databases and links (USA National Institute of Standards and Technology)
NIST TREC Video Retrieval Evaluation Database (USA National Institute of Standards and Technology)
Princeton Shape Benchmark (Princeton Shape Retrieval and Analysis Group)
Queensland cross media dataset - millions of images and text documents for "cross-media" retrieval (Yi Yang)
TOSCA 3D shape database (Bronstein, Bronstein, Kimmel)

Object Databases

2.5D/3D Datasets of various objects and scenes (Ajmal Mian)
Amsterdam Library of Object Images (ALOI): 100K views of 1K objects (University of Amsterdam/Intelligent Sensory Information Systems)
Caltech 101 (now 256) category object recognition database (Li Fei-Fei, Marco Andreeto, Marc'Aurelio Ranzato)
Columbia COIL-100 3D object multiple views (Columbia University)
Densely sampled object views: 2500 views of 2 objects, eg for view-based recognition and modeling (Gabriele Peters, Universiteit Dortmund)
German Traffic Sign Detection Benchmark (Ruhr-Universitat Bochum)
GRAZ-02 Database (Bikes, cars, people) (A. Pinz)
Linkoping 3D Object Pose Estimation Database (Fredrik Viksten and Per-Erik Forssen)
Microsoft Object Class Recognition image databases (Antonio Criminisi, Pushmeet Kohli, Tom Minka, Carsten Rother, Toby Sharp, Jamie Shotton, John Winn)
Microsoft salient object databases (labeled by bounding boxes) (Liu, Sun Zheng, Tang, Shum)
MIT CBCL Car Data (Center for Biological and Computational Learning)
MIT CBCL StreetScenes Challenge Framework: (Stan Bileschi)
NEC Toy animal object recognition or categorization database (Hossein Mobahi)
NORB 50 toy image database (NYU)
PASCAL Image Database (motorbikes, cars, cows) (PASCAL Consortium)
PASCAL 2007 Challange Image Database (motorbikes, cars, cows) (PASCAL Consortium)
PASCAL 2008 Challange Image Database (PASCAL Consortium)
PASCAL 2009 Challange Image Database (PASCAL Consortium)
PASCAL 2010 Challange Image Database (PASCAL Consortium)
PASCAL 2011 Challange Image Database (PASCAL Consortium)
PASCAL 2012 Challange Image Database Category classification, detection, and segmentation, and still-image action classification (PASCAL Consortium)
UIUC Car Image Database (UIUC)
UIUC Dataset of 3D object categories (S. Savarese and L. Fei-Fei)
Venezia 3D object-in-clutter recognition and segmentation (Emanuele Rodola)

People, Pedestrian, Eye/Iris, Template Detection/Tracking Databases

3D KINECT Gender Walking data base (L. Igual, A. Lapedriza, R. Borràs from UB, CVC and UOC, Spain)
Caltech Pedestrian Dataset (P. Dollar, C. Wojek, B. Schiele and P. Perona)
CASIA gait database (Chinese Academy of Sciences)
CASIA-IrisV3 (Chinese Academy of Sciences, T. N. Tan, Z. Sun)
CAVIAR project video sequences with tracking and behavior ground truth (CAVIAR team/Edinburgh University - EC project IST-2001-37540)
Daimler Pedestrian Detection Benchmark 21790 images with 56492 pedestrians plus empty scenes (M. Enzweiler, D. M. Gavrila)
Driver Monitoring Video Dataset (RobeSafe + Jesus Nuevo-Chiquero)
Edinburgh overhead camera person tracking dataset (Bob Fisher, Bashia Majecka, Gurkirt Singh, Rowland Sillito)
Eyetracking database summary (Stefan Winkler)
HAT database of 27 human attributes (Gaurav Sharma, Frederic Jurie)
INRIA Person Dataset (Navneet Dalal)
ISMAR09 ground truth video dataset for template-based (i.e. planar) tracking algorithms (Sebastian Lieberknecht)
MIT CBCL Pedestrian Data (Center for Biological and Computational Learning)
MIT eye tracking database (1003 images) (Judd et al)
Notre Dame Iris Image Dataset (Patrick J. Flynn)
PETS 2009 Crowd Challange dataset (Reading University & James Ferryman)
PETS: Performance Evaluation of Tracking and Surveillance (Reading University & James Ferryman)
PETS Winter 2009 workshop data (Reading University & James Ferryman)
UBIRIS: Noisy Visible Wavelength Iris Image Databases (University of Beira)
Univ of Central Florida - Crowd Dataset (Saad Ali)
Univ of Central Florida - Crowd Flow Segmentation datasets (Saad Ali)
York Univ Eye Tracking Dataset (120 images) (Neil Bruce)

Segmentation

Alpert et al. Segmentation evaluation database (Sharon Alpert, Meirav Galun, Ronen Basri, Achi Brandt)
Berkeley Segmentation Dataset and Benchmark (David Martin and Charless Fowlkes)
GrabCut Image database (C. Rother, V. Kolmogorov, A. Blake, M. Brown)
LabelMe images database and online annotation tool (Bryan Russell, Antonio Torralba, Kevin Murphy, William Freeman)

Surveillance

AVSS07: Advanced Video and Signal based Surveillance 2007 datasets (Andrea Cavallaro)
ETISEO Video Surveillance Download Datasets (INRIA Orion Team and others)
Heriot Watt Summary of datasets for human tracking and surveillance (Zsolt Husz)
SPEVI: Surveillance Performance EValuation Initiative (Queen Mary University London)
Udine Trajectory-based anomalous event detection dataset - synthetic trajectory datasets with outliers (Univ of Udine Artificial Vision and Real Time Systems Laboratory)

Textures

Color texture images by category (textures.forrest.cz)
Columbia-Utrecht Reflectance and Texture Database (Columbia & Utrecht Universities)
DynTex: Dynamic texture database (Renaud Piteri, Mark Huiskes and Sandor Fazekas)
Oulu Texture Database (Oulu University)
Prague Texture Segmentation Data Generator and Benchmark (Mikes, Haindl)
Uppsala texture dataset of surfaces and materials - fabrics, grains, etc.
Vision Texture (MIT Media Lab)

General Videos

Large scale YouTube video dataset - 156,823 videos (2,907,447 keyframes) crawled from YouTube videos (Yi Yang)

Other Collections

CANTATA Video and Image Database Index site (Multitel)
Computer Vision Homepage list of test image databases (Carnegie Mellon Univ)
ETHZ various, including 3D head pose, shape classes, pedestrians, pedestrians, buildings (ETH Zurich, Computer Vision Lab)
Leibe's Collection of people/vehicle/object databases (Bastian Leibe)
Lotus Hill Image Database Collection with Ground Truth (Sealeen Ren, Benjamin Yao, Michael Yang)
Oxford Misc, including Buffy, Flowers, TV characters, Buildings, etc (Oxford Visual geometry Group)
PEIPA Image Database Summary (Pilot European Image Processing Archive)
Univ of Bern databases on handwriting, online documents, string edit and graph matching (Univ of Bern, Computer Vision and Artificial Intelligence)
USC Annotated Computer Vision Bibliography database publication summary (Keith Price)
USC-SIPI image databases: texture, aerial, favorites (eg. Lena) (USC Signal and Image Processing Institute)

Miscellaneous

3D mesh watermarking benchmark dataset (Guillaume Lavoue)
Active Appearance Models datasets (Mikkel B. Stegmann)
Aircraft tracking (Ajmal Mian)
Cambridge Motion-based Segmentation and Recognition Dataset (Brostow, Shotton, Fauqueur, Cipolla)
Catadioptric camera calibration images (Yalin Bastanlar)
Chars74K dataset - 74 English and Kannada characters (Teo de Campos - [email protected])
COLD (COsy Localization Database) - place localization (Ullah, Pronobis, Caputo, Luo, and Jensfelt)
Columbia Camera Response Functions: Database (DoRF) and Model (EMOR) (M.D. Grossberg and S.K. Nayar)
Columbia Database of Contaminants' Patterns and Scattering Parameters (Jinwei Gu, Ravi Ramamoorthi, Peter Belhumeur, Shree Nayar)
Dense outdoor correspondence ground truth datasets, for optical flow and local keypoint evaluation (Christoph Strecha)
DTU controlled motion and lighting image dataset (135K images) (Henrik Aanaes)
EISATS: .enpeda.. Image Sequence Analysis Test Site (Auckland University Multimedia Imaging Group)
FlickrLogos-32 - 8240 images of 32 product logos (Stefan Romberg)
Flowchart images (Allan Hanbury)
Geometric Context - scene interpretation images (Derek Hoiem)
Image/video quality assessment database summary (Stefan Winkler)
INRIA feature detector evaluation sequences (Krystian Mikolajczyk)
INRIA's PERCEPTION's database of images and videos gathered with several synchronized and calibrated cameras (INRIA Rhone-Alpes)
INRIA's Synchronized and calibrated binocular/binaural data sets with head movements (INRIA Rhone-Alpes)
KITTI dataset for stereo, optical flow and visual odometry (Geiger, Lenz, Urtasun)
Large scale 3D point cloud data from terrestrial LiDAR scanning (Andreas Nuechter)
Linkoping Rolling Shutter Rectification Dataset (Per-Erik Forssen and Erik Ringaby)
Middlebury College stereo vision research datasets (Daniel Scharstein and Richard Szeliski)
MPI-Sintel optical flow evaluation dataset (Michael Black)
Multiview stereo images with laser based groundtruth (ESAT-PSI/VISICS,FGAN-FOM,EPFL/IC/ISIM/CVLab)
The Cancer Imaging Archive (National Cancer Institute)
NCI Cancer Image Archive - prostate images (National Cancer Institute)
NIST 3D Interest Point Detection (Helin Dutagaci, Afzal Godil)
NRCS natural resource/agricultural image database (USDA Natural Resources Conservation Service)
Occlusion detection test data (Andrew Stein)
The Open Video Project (Gary Marchionini, Barbara M. Wildemuth, Gary Geisler, Yaxiao Song)
Pics 'n' Trails - Dataset of Continuously archived GPS and digital photos (Gamhewage Chaminda de Silva)
PRINTART: Artistic images of prints of well known paintings, including detail annotations. A benchmark for automatic annotation and retrieval tasks with this database was published at ECCV. (Nuno Miguel Pinho da Silva)
RAWSEEDS SLAM benchmark datasets (Rawseeds Project)
Robotic 3D Scan Repository - 3D point clouds from robotic experiments of scenes (Osnabruck and Jacobs Universities)
ROMA (ROad MArkings) : Image database for the evaluation of road markings extraction algorithms (Jean-Philippe Tarel, et al)
Stuttgart Range Image Database - 66 views of 45 objects
UCL Ground Truth Optical Flow Dataset (Oisin Mac Aodha)
Univ of Genoa Datasets for disparity and optic flow evaluation (Manuela Chessa)
Validation and Verification of Neural Network Systems (Francesco Vivarelli)
VSD: Technicolor Violent Scenes Dataset - a collection of ground-truth files based on the extraction of violent events in movies

WILD: Weather and Illumunation Database (S. Narasimhan, C. Wang. S. Nayar, D. Stolyarov, K. Garg, Y. Schechner, H. Peri)

转自博客：http://blog.csdn.net/tiandijun

你可能感兴趣的:(常用图像数据集大全（分类，跟踪，分割，检测等）)

机器学习与深度学习间关系与区别 ℒℴѵℯ心·动ꦿ໊ོ꫞ 人工智能学习深度学习 python
一、机器学习概述定义机器学习（MachineLearning,ML）是一种通过数据驱动的方法，利用统计学和计算算法来训练模型，使计算机能够从数据中学习并自动进行预测或决策。机器学习通过分析大量数据样本，识别其中的模式和规律，从而对新的数据进行判断。其核心在于通过训练过程，让模型不断优化和提升其预测准确性。主要类型1.监督学习（SupervisedLearning）监督学习是指在训练数据集中包含输入
OC语言多界面传值五大方式 Magnetic_h ios ui 学习 objective-c 开发语言
前言在完成暑假仿写项目时，遇到了许多需要用到多界面传值的地方，这篇博客来总结一下比较常用的五种多界面传值的方式。属性传值属性传值一般用前一个界面向后一个界面传值，简单地说就是通过访问后一个视图控制器的属性来为它赋值，通过这个属性来做到从前一个界面向后一个界面传值。首先在后一个界面中定义属性@interfaceBViewController:UIViewController@propertyNSSt
微服务下功能权限与数据权限的设计与实现 nbsaas-boot 微服务 java 架构
在微服务架构下，系统的功能权限和数据权限控制显得尤为重要。随着系统规模的扩大和微服务数量的增加，如何保证不同用户和服务之间的访问权限准确、细粒度地控制，成为设计安全策略的关键。本文将讨论如何在微服务体系中设计和实现功能权限与数据权限控制。1.功能权限与数据权限的定义功能权限：指用户或系统角色对特定功能的访问权限。通常是某个用户角色能否执行某个操作，比如查看订单、创建订单、修改用户资料等。数据权限：
理解Gunicorn：Python WSGI服务器的基石范范0825 ipython linux 运维
理解Gunicorn：PythonWSGI服务器的基石介绍Gunicorn，全称GreenUnicorn，是一个为PythonWSGI（WebServerGatewayInterface）应用设计的高效、轻量级HTTP服务器。作为PythonWeb应用部署的常用工具，Gunicorn以其高性能和易用性著称。本文将介绍Gunicorn的基本概念、安装和配置，帮助初学者快速上手。1.什么是Gunico
学点心理知识，呵护孩子健康静候花开_7090
昨天听了华中师范大学教育管理学系副教授张玲老师的《哪里才是学生心理健康的最后庇护所，超越教育与技术的思考》的讲座。今天又重新学习了一遍，收获匪浅。张玲博士也注意到了当今社会上的孩子由于心理问题导致的自残、自杀及伤害他人等恶性事件。她向我们普及了一个重要的命题，她说心理健康的一些基本命题，我们与我们通常的一些教育命题是不同的，她还举了几个例子，让我们明白我们原来以为的健康并非心理学上的健康。比如如果
扫地机类清洁产品之直流无刷电机控制悟空胆好小清洁服务机器人单片机人工智能
扫地机类清洁产品之直流无刷电机控制1.1前言扫地机产品有很多的电机控制，滚刷电机1个，边刷电机1-2个，清水泵电机，风机一个，部分中高端产品支持抹布功能，也就是存在抹布盘电机，还有追觅科沃斯石头等边刷抬升电机，滚刷抬升电机等的，这些电机有直流有刷电机，直接无刷电机，步进电机，电磁阀，挪动泵等不同类型。电机的原理，驱动控制方式也不行。接下来一段时间的几个文章会作个专题分析分享。直流有刷电机会自动持续
Linux下QT开发的动态库界面弹出操作（SDL2） 13jjyao QT类 qt 开发语言 sdl2 linux
需求：操作系统为linux，开发框架为qt，做成需带界面的qt动态库，调用方为java等非qt程序难点：调用方为java等非qt程序，也就是说调用方肯定不带QApplication::exec()，缺少了这个，QTimer等事件和QT创建的窗口将不能弹出(包括opencv也是不能弹出)；这与qt调用本身qt库是有本质的区别的思路：1.调用方缺QApplication::exec()，那么我们在接口
向内而求陈陈_19b4
10月27日，阴。阅读书目:《次第花开》。作者:希阿荣博堪布，是当今藏传佛家宁玛派最伟大的上师法王，如意宝晋美彭措仁波切颇具影响力的弟子之一。多年以来，赴海内外各地弘扬佛法，以正式授课、现场开示、发表文章等多种方法指导佛学弟子修行佛法。代表作《寂静之道》、《生命这出戏》、《透过佛法看世界》自出版以来一直是佛教类书籍中的畅销书。图片发自App金句:1.佛陀说，一切痛苦的根源在于我们长期以来对自身及外
消息中间件有哪些常见类型 xmh-sxh-1314 java
消息中间件根据其设计理念和用途，可以大致分为以下几种常见类型：点对点消息队列（Point-to-PointMessagingQueues）：在这种模型中，消息被发送到特定的队列中，消费者从队列中取出并处理消息。队列中的消息只能被一个消费者消费，消费后即被删除。常见的实现包括IBM的MQSeries、RabbitMQ的部分使用场景等。适用于任务分发、负载均衡等场景。发布/订阅消息模型（Pub/Sub
ArcGIS栅格计算器常见公式（赋值、0和空值的转换、补充栅格空值）研学随笔 arcgis 经验分享
我们在使用ArcGIS时通常经常用到栅格计算器，今天主要给大家介绍我日常中经常用到的几个公式，供大家参考学习。将特定值（-9999）赋值为0，例如-9999.Con("raster"==-9999,0,"raster")2.给空值赋予特定的值（如0）Con(IsNull("raster"),0,"raster")3.将特定的栅格值(如1)赋值为空值，其他保留原值SetNull("raster"==
《大清方方案》| 第二话谁佐清欢
和珅究竟说了些什么？竟能令堂堂九五之尊龙颜失色！此处暂且按下不表；单说这位乾隆皇帝，果真不愧是康熙从小带过的，一旦决定了要做的事，便杀伐决断毫不含糊。他当即亲自拟旨，着令和珅为钦差大臣，全权负责处理方方事件，并钦赐尚方宝剑，遇急则三品以下官员可先斩后奏。和珅身负皇上重托，岂敢有半点怠慢，当夜即率领相关人等，马不停蹄杀奔江汉。这一路上，和珅的几位幕僚一直在商讨方方事件的处置方案。有位年轻幕僚建议快刀
Python数据分析与可视化实战指南 William数据分析 python python 数据
在数据驱动的时代，Python因其简洁的语法、强大的库生态系统以及活跃的社区，成为了数据分析与可视化的首选语言。本文将通过一个详细的案例，带领大家学习如何使用Python进行数据分析，并通过可视化来直观呈现分析结果。一、环境准备1.1安装必要库在开始数据分析和可视化之前，我们需要安装一些常用的库。主要包括pandas、numpy、matplotlib和seaborn等。这些库分别用于数据处理、数学
C#中使用split分割字符串互联网打工人no1 c#
1、用字符串分隔：usingSystem.Text.RegularExpressions;stringstr="aaajsbbbjsccc";string[]sArray=Regex.Split(str,"js",RegexOptions.IgnoreCase);foreach(stringiinsArray)Response.Write(i.ToString()+"");输出结果：aaabbbc
git常用命令笔记咩酱-小羊 git 笔记
###用习惯了idea总是不记得git的一些常见命令，需要用到的时候总是担心旁边站了人~~~记个笔记@_@，告诉自己看笔记不丢人初始化初始化一个新的Git仓库gitinit配置配置用户信息gitconfig--globaluser.name"YourName"gitconfig--globaluser.email"[email protected]"基本操作克隆远程仓库gitclone查看
怎么起诉借钱不还的人？怎样起诉欠款不还的人？影子爱学习
怎么起诉借钱不还的人？怎样起诉欠款不还的人？如果遇到难以解决的法律问题，我们可以匹配专业律师。例如：婚姻家庭（离婚纠纷）、刑事辩护、合同纠纷、债权债务、房产（继承）纠纷、交通事故、劳动争议、人身损害、公司相关法律事务（法律顾问）等咨询推荐手机/微信:15633770876【全国案件皆可】借钱不还起诉对方需要哪些资料起诉欠钱不还的，一般需要的材料包括以下这些：借据、收据、欠条、付款凭证等证据，以及向
相信相信的力量孙丽_cdb3
孙丽中级十期坚持分享第345天有一个特别有哲理的故事：有一只老鹰下了蛋，这个蛋，不知怎的就滚到了鸡窝里去了，鸡也下了一窝蛋，然后鸡妈妈把这些蛋全都浮出来了，孵出来之后等小鸡长大一点了，就觉得鹰蛋孵出来的那只小鹰怪模怪样，这些小鸡都嘲笑它，真难看，真笨，丑死了，那只小鹰觉得自己真是谁也不像，真是不好看，后来鸡妈妈也不喜欢他，我怎么生出你这样的孩子来了？真烦人，后来这群小鸡和小鹰一起生活，有一天，老鹰
LLM 词汇表落难Coder LLMs NLP 大语言模型大模型 llama 人工智能
Contextwindow“上下文窗口”是指语言模型在生成新文本时能够回溯和参考的文本量。这不同于语言模型训练时所使用的大量数据集，而是代表了模型的“工作记忆”。较大的上下文窗口可以让模型理解和响应更复杂和更长的提示，而较小的上下文窗口可能会限制模型处理较长提示或在长时间对话中保持连贯性的能力。Fine-tuning微调是使用额外的数据进一步训练预训练语言模型的过程。这使得模型开始表示和模仿微调数
基于社交网络算法优化的二维最大熵图像分割智能算法研学社（Jack旭）智能优化算法应用图像分割算法 php 开发语言
智能优化算法应用：基于社交网络优化的二维最大熵图像阈值分割-附代码文章目录智能优化算法应用：基于社交网络优化的二维最大熵图像阈值分割-附代码1.前言2.二维最大熵阈值分割原理3.基于社交网络优化的多阈值分割4.算法结果：5.参考文献：6.Matlab代码摘要：本文介绍基于最大熵的图像分割，并且应用社交网络算法进行阈值寻优。1.前言阅读此文章前，请阅读《图像分割：直方图区域划分及信息统计介绍》htt
509. 斐波那契数(每日一题) lzyprime
lzyprime博客(github)创建时间：2021.01.04qq及邮箱：2383518170leetcode笔记题目描述斐波那契数，通常用F(n)表示，形成的序列称为斐波那契数列。该数列由0和1开始，后面的每一项数字都是前面两项数字的和。也就是：F(0)=0，F(1)=1F(n)=F(n-1)+F(n-2)，其中n>1给你n，请计算F(n)。示例1：输入：2输出：1解释：F(2)=F(1)+
Git常用命令－修改远程仓库地址猿大师 Linux Java git java
查看远程仓库地址gitremote-v返回结果originhttps://git.coding.net/＊＊＊＊＊.git(fetch)originhttps://git.coding.net/＊＊＊＊＊.git(push)修改远程仓库地址gitremoteset-urloriginhttps://git.coding.net/＊＊＊＊＊.git先删除后增加远程仓库地址gitremotermori
使用LLaVa和Ollama实现多模态RAG示例 llzwxh888 python 人工智能开发语言
本文将详细介绍如何使用LLaVa和Ollama实现多模态RAG（检索增强生成），通过提取图像中的结构化数据、生成图像字幕等功能来展示这一技术的强大之处。安装环境首先，您需要安装以下依赖包：!pipinstallllama-index-multi-modal-llms-ollama!pipinstallllama-index-readers-file!pipinstallunstructured!p
GitHub上克隆项目 bigbig猩猩 github
从GitHub上克隆项目是一个简单且直接的过程，它允许你将远程仓库中的项目复制到你的本地计算机上，以便进行进一步的开发、测试或学习。以下是一个详细的步骤指南，帮助你从GitHub上克隆项目。一、准备工作1.安装Git在克隆GitHub项目之前，你需要在你的计算机上安装Git工具。Git是一个开源的分布式版本控制系统，用于跟踪和管理代码变更。你可以从Git的官方网站（https://git-scm.
春季养肝正当时 dxn悟
重温快乐2023年2月4日立春。春天来了，春暖花开，小鸟欢唱，那在这样的季节我们如何养肝呢？自然界的春季对应中医五行的木，人体五脏肝属木，“木曰曲直”，是以树干曲曲直直地向上、向外伸长舒展的生发姿态，来形容具有生长、升发、条达、舒畅等特征的食物及现象。根据中医天人相应的理念，肝五行属木，喜条达，主疏泄，与春天相应，所以春天最适合养肝。养肝首先要少生气，因为肝喜条达恶抑郁。人体五志肝为怒，生气发怒最
关于城市旅游的HTML网页设计——(旅游风景云南 5页)HTML+CSS+JavaScript 二挡起步 web前端期末大作业 javascript html css 旅游风景
⛵源码获取文末联系✈Web前端开发技术描述网页设计题材，DIV+CSS布局制作,HTML+CSS网页设计期末课程大作业|游景点介绍|旅游风景区|家乡介绍|等网站的设计与制作|HTML期末大学生网页设计作业，Web大学生网页HTML：结构CSS：样式在操作方面上运用了html5和css3，采用了div+css结构、表单、超链接、浮动、绝对定位、相对定位、字体样式、引用视频等基础知识JavaScrip
HTML网页设计制作大作业（div+css）云南我的家乡旅游景点带文字滚动二挡起步 web前端期末大作业 web设计网页规划与设计 html css javascript dreamweaver 前端
Web前端开发技术描述网页设计题材，DIV+CSS布局制作,HTML+CSS网页设计期末课程大作业游景点介绍|旅游风景区|家乡介绍|等网站的设计与制作HTML期末大学生网页设计作业HTML：结构CSS：样式在操作方面上运用了html5和css3，采用了div+css结构、表单、超链接、浮动、绝对定位、相对定位、字体样式、引用视频等基础知识JavaScript：做与用户的交互行为文章目录前端学习路线
Day17笔记-高阶函数 ~在杰难逃~ Python 笔记 python 开发语言 pycharm 数据分析
高阶函数【重点掌握】函数的本质：函数是一个变量，函数名是一个变量名，一个函数可以作为另一个函数的参数或返回值使用如果A函数作为B函数的参数，B函数调用完成之后，会得到一个结果，则B函数被称为高阶函数常用的高阶函数：map(),reduce(),filter(),sorted()1.map()map(func,iterable)，返回值是一个iterator【容器，迭代器】func:函数iterab
【目标检测数据集】卡车数据集1073张VOC+YOLO格式熬夜写代码的平头哥∰ 目标检测 YOLO 人工智能
数据集格式：PascalVOC格式+YOLO格式(不包含分割路径的txt文件，仅仅包含jpg图片以及对应的VOC格式xml文件和yolo格式txt文件)图片数量(jpg文件个数)：1073标注数量(xml文件个数)：1073标注数量(txt文件个数)：1073标注类别数：1标注类别名称:["truck"]每个类别标注的框数：truck框数=1120总框数：1120使用标注工具：labelImg标注
人工智能时代，程序员如何保持核心竞争力？ jmoych 人工智能
随着AIGC（如chatgpt、midjourney、claude等）大语言模型接二连三的涌现，AI辅助编程工具日益普及，程序员的工作方式正在发生深刻变革。有人担心AI可能取代部分编程工作，也有人认为AI是提高效率的得力助手。面对这一趋势,程序员应该如何应对?是专注于某个领域深耕细作，还是广泛学习以适应快速变化的技术环境?又或者，我们是否应该将重点转向AI无法轻易替代的软技能？让我们一起探讨程序员
每日算法&面试题，大厂特训二十八天——第二十天（树）肥学 ⚡算法题⚡面试题每日精进 java 算法数据结构
目录标题导读算法特训二十八天面试题点击直接资料领取导读肥友们为了更好的去帮助新同学适应算法和面试题，最近我们开始进行专项突击一步一步来。上一期我们完成了动态规划二十一天现在我们进行下一项对各类算法进行二十八天的一个小总结。还在等什么快来一起肥学进行二十八天挑战吧！！特别介绍小白练手专栏，适合刚入手的新人欢迎订阅编程小白进阶python有趣练手项目里面包括了像《机器人尬聊》《恶搞程序》这样的有趣文章
2022现在哪个打车软件比较好用又便宜实惠的打车软件合集高省APP珊珊
这是一个信息高速传播的社会。信息可以通过手机，微信，自媒体，抖音等方式进行传播。但同时这也是一个交通四通发达的社会。高省APP，是2022年推出的平台，0投资，0风险、高省APP佣金更高，模式更好，终端用户不流失。【高省】是一个自用省钱佣金高，分享推广赚钱多的平台，百度有几百万篇报道，也期待你的加入。珊珊导师，高省邀请码777777，注册送2皇冠会员，送万元推广大礼包，教你如何1年做到百万团队。高
Algorithm 香水浓 java Algorithm
冒泡排序 public static void sort(Integer[] param) { for (int i = param.length - 1; i > 0; i--) { for (int j = 0; j < i; j++) { int current = param[j]; int next = param[j + 1];
mongoDB 复杂查询表达式开窍的石头 mongodb
1:count Pg: db.user.find().count(); 统计多少条数据 2:不等于$ne Pg: db.user.find({_id:{$ne:3}},{name:1,sex:1,_id:0}); 查询id不等于3的数据。 3：大于$gt $gte(大于等于) &n
Jboss Java heap space异常解决方法, jboss OutOfMemoryError : PermGen space 0624chenhong jvm jboss
转自 http://blog.csdn.net/zou274/article/details/5552630 解决办法： window->preferences->java->installed jres->edit jre 把default vm arguments 的参数设为-Xms64m -Xmx512m ----------------
文件上传下载解析相对路径不懂事的小屁孩文件上传
有点坑吧，弄这么一个简单的东西弄了一天多，身边还有大神指导着，网上各种百度着。下面总结一下遇到的问题：文件上传，在页面上传的时候，不要想着去操作绝对路径，浏览器会对客户端的信息进行保护，避免用户信息收到攻击。在上传图片，或者文件时，使用form表单来操作。前台通过form表单传输一个流到后台，而不是ajax传递参数到后台，代码如下: <form action=&
怎么实现qq空间批量点赞换个号韩国红果果 qq
纯粹为了好玩！！逻辑很简单 1 打开浏览器console；输入以下代码。先上添加赞的代码 var tools={}; //添加所有赞 function init(){ document.body.scrollTop=10000; setTimeout(function(){document.body.scrollTop=0;},2000);//加
判断是否为中文灵静志远中文
方法一： public class Zhidao { public static void main(String args[]) { String s = "sdf灭礌 kjl d{';\fdsjlk是"; int n=0; for(int i=0; i<s.length(); i++) { n = (int)s.charAt(i); if((
一个电话面试后总结 a-john 面试
今天，接了一个电话面试，对于还是初学者的我来说，紧张了半天。面试的问题分了层次，对于一类问题，由简到难。自己觉得回答不好的地方作了一下总结：在谈到集合类的时候，举几个常用的集合类，想都没想，直接说了list,map。然后对list和map分别举几个类型： list方面：ArrayList,LinkedList。在谈到他们的区别时，愣住了
MSSQL中Escape转义的使用 aijuans MSSQL
IF OBJECT_ID('tempdb..#ABC') is not null drop table tempdb..#ABC create table #ABC ( PATHNAME NVARCHAR(50) ) insert into #ABC SELECT N'/ABCDEFGHI' UNION ALL SELECT N'/ABCDGAFGASASSDFA' UNION ALL
一个简单的存储过程 asialee mysql 存储过程构造数据批量插入
今天要批量的生成一批测试数据，其中中间有部分数据是变化的，本来想写个程序来生成的，后来想到存储过程就可以搞定，所以随手写了一个，记录在此： DELIMITER $$ DROP PROCEDURE IF EXISTS inse
annot convert from HomeFragment_1 to Fragment 百合不是茶 android 导包错误
创建了几个类继承Fragment, 需要将创建的类存储在ArrayList<Fragment>中; 出现不能将new 出来的对象放到队列中,原因很简单; 创建类时引入包是:import android.app.Fragment; 创建队列和对象时使用的包是:import android.support.v4.ap
Weblogic10两种修改端口的方法 bijian1013 weblogic 端口号配置管理 config.xml
一.进入控制台进行修改 1.进入控制台: http://127.0.0.1:7001/console 2.展开左边树菜单域结构->环境->服务器-->点击AdminServer(管理) &
mysql 操作指令征客丶 mysql
一、连接mysql 进入 mysql 的安装目录； $ bin/mysql -p [host IP 如果是登录本地的mysql 可以不写 -p 直接 -u] -u [userName] -p 输入密码，回车，接连；二、权限操作［如果你很了解mysql数据库后，你可以直接去修改系统表，然后用 mysql> flush privileges; 指令让权限生效］ 1、赋权 mys
【Hive一】Hive入门 bit1129 hive
Hive安装与配置 Hive的运行需要依赖于Hadoop，因此需要首先安装Hadoop2.5.2，并且Hive的启动前需要首先启动Hadoop。 Hive安装和配置的步骤 1. 从如下地址下载Hive0.14.0 http://mirror.bit.edu.cn/apache/hive/ 2.解压hive，在系统变
ajax 三种提交请求的方法 BlueSkator Ajax jqery
1、ajax 提交请求 $.ajax({ type:"post", url : "${ctx}/front/Hotel/getAllHotelByAjax.do", dataType : "json", success : function(result) { try { for(v
mongodb开发环境下的搭建入门 braveCS 运维
linux下安装mongodb 1）官网下载mongodb-linux-x86_64-rhel62-3.0.4.gz 2）linux 解压 gzip -d mongodb-linux-x86_64-rhel62-3.0.4.gz; mv mongodb-linux-x86_64-rhel62-3.0.4 mongodb-linux-x86_64-rhel62-
编程之美-最短摘要的生成 bylijinnan java 数据结构算法编程之美
import java.util.HashMap; import java.util.Map; import java.util.Map.Entry; public class ShortestAbstract { /** * 编程之美最短摘要的生成 * 扫描过程始终保持一个[pBegin,pEnd]的range,初始化确保[pBegin,pEnd]的ran
json数据解析及typeof chengxuyuancsdn js typeof json解析
// json格式 var people='{"authors": [{"firstName": "AAA","lastName": "BBB"},' +' {"firstName": "CCC&
流程系统设计的层次和目标 comsci 设计模式数据结构 sql 框架脚本
流程系统设计的层次和目标
RMAN List和report 命令 daizj oracle list report rman
LIST 命令使用RMAN LIST 命令显示有关资料档案库中记录的备份集、代理副本和映像副本的信息。使用此命令可列出： • RMAN 资料档案库中状态不是AVAILABLE 的备份和副本 • 可用的且可以用于还原操作的数据文件备份和副本 • 备份集和副本，其中包含指定数据文件列表或指定表空间的备份 • 包含指定名称或范围的所有归档日志备份的备份集和副本 • 由标记、完成时间、可
二叉树:红黑树 dieslrae 二叉树
红黑树是一种自平衡的二叉树,它的查找,插入,删除操作时间复杂度皆为O(logN),不会出现普通二叉搜索树在最差情况时时间复杂度会变为O(N)的问题. 红黑树必须遵循红黑规则,规则如下 1、每个节点不是红就是黑。 2、根总是黑的 &
C语言homework3，7个小题目的代码 dcj3sjt126com c
1、打印100以内的所有奇数。 # include <stdio.h> int main(void) { int i; for (i=1; i<=100; i++) { if (i%2 != 0) printf("%d ", i); } return 0; } 2、从键盘上输入10个整数，
自定义按钮, 图片在上, 文字在下, 居中显示 dcj3sjt126com 自定义
#import <UIKit/UIKit.h> @interface MyButton : UIButton -(void)setFrame:(CGRect)frame ImageName:(NSString*)imageName Target:(id)target Action:(SEL)action Title:(NSString*)title Font:(CGFloa
MySQL查询语句练习题，测试足够用了 flyvszhb sql mysql
http://blog.sina.com.cn/s/blog_767d65530101861c.html 1.创建student和score表 CREATE TABLE student ( id INT(10) NOT NULL UNIQUE PRIMARY KEY , name VARCHAR
转：MyBatis Generator 详解 happyqing mybatis
MyBatis Generator 详解 http://blog.csdn.net/isea533/article/details/42102297 MyBatis Generator详解 http://git.oschina.net/free/Mybatis_Utils/blob/master/MybatisGeneator/MybatisGeneator.
让程序员少走弯路的14个忠告 jingjing0907 工作计划学习
无论是谁，在刚进入某个领域之时，有再大的雄心壮志也敌不过眼前的迷茫：不知道应该怎么做，不知道应该做什么。下面是一名软件开发人员所学到的经验，希望能对大家有所帮助 1.不要害怕在工作中学习。只要有电脑，就可以通过电子阅读器阅读报纸和大多数书籍。如果你只是做好自己的本职工作以及分配的任务，那是学不到很多东西的。如果你盲目地要求更多的工作，也是不可能提升自己的。放
nginx和NetScaler区别流浪鱼 nginx
NetScaler是一个完整的包含操作系统和应用交付功能的产品，Nginx并不包含操作系统，在处理连接方面，需要依赖于操作系统，所以在并发连接数方面和防DoS攻击方面，Nginx不具备优势。 2.易用性方面差别也比较大。Nginx对管理员的水平要求比较高，参数比较多，不确定性给运营带来隐患。在NetScaler常见的配置如健康检查，HA等，在Nginx上的配置的实现相对复杂。 3.策略灵活度方
第11章动画效果（下） onestopweb 动画
index.html <!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Transitional//EN" "http://www.w3.org/TR/xhtml1/DTD/xhtml1-transitional.dtd"> <html xmlns="http://www.w3.org/
FAQ - SAP BW BO roadmap blueoxygen BO BW
http://www.sdn.sap.com/irj/boc/business-objects-for-sap-faq Besides, I care that how to integrate tightly. By the way, for BW consultants, please just focus on Query Designer which i
关于java堆内存溢出的几种情况 tomcat_oracle java jvm jdk thread
【情况一】：　　 java.lang.OutOfMemoryError: Java heap space：这种是java堆内存不够，一个原因是真不够，另一个原因是程序中有死循环；　　如果是java堆内存不够的话，可以通过调整JVM下面的配置来解决：　　<jvm-arg>-Xms3062m</jvm-arg> 　　<jvm-arg>-Xmx
Manifest.permission_group权限组阿尔萨斯 Permission
结构继承关系 public static final class Manifest.permission_group extends Object java.lang.Object android. Manifest.permission_group 常量 ACCOUNTS 直接通过统计管理器访问管理的统计 COST_MONEY可以用来让用户花钱但不需要通过与他们直接牵涉的权限 D