Part 3 of 7 · Face Recognition App

How Face Recognition Actually Works, in Plain Words

Recognition sounds like magic until you see the mechanism, and the mechanism is beautifully plain, every face becomes a list of 128 numbers, and recognising someone is measuring how close two lists are. This post is that mechanism, explained through the real matching function at the heart of my app, because once you understand the encoding, everything about the system’s behaviour, including its failures, makes sense.

The 128 numbers are called an encoding, produced by a neural network trained on millions of faces to output similar numbers for the same person and distant numbers for different people, across lighting, angle, and expression. The training is the research-grade part living in dlib, using it is three function calls, and here is the actual recognition path from my app, processing one camera frame:

import cv2, face_recognition

def recognize_face(self, frame):
    rgb_frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
    face_locations = face_recognition.face_locations(rgb_frame)
    face_encodings = face_recognition.face_encodings(rgb_frame, face_locations)

    recognized = []
    for face_encoding, face_location in zip(face_encodings, face_locations):
        matches = face_recognition.compare_faces(
            list(self.known_faces.values()), face_encoding)
        name = 'Unknown'
        if True in matches:
            match_index = matches.index(True)
            name = list(self.known_faces.keys())[match_index]
        recognized.append((name, face_location))
    return recognized

Walk the stages. The colour conversion first, because OpenCV delivers frames in BGR order and face_recognition expects RGB, and skipping cv2.cvtColor is the classic silent degradation, everything runs, accuracy quietly suffers. face_locations finds where faces are in the frame, face_encodings turns each found face into its 128 numbers, computed only at those locations, and compare_faces measures each against the whole known set, returning a list of booleans, is this person close enough to each known encoding. A True means the distance fell under the tolerance, and the name comes from the matching entry, no match anywhere means Unknown, which the alerts post turns into a feature.

Understanding distance-under-tolerance explains the system’s whole character. The default tolerance is 0.6, and it is a dial between error types, tighter rejects more strangers and more bad photos of friends, looser accepts more friends and more strangers, there is no setting without a trade, only a trade chosen consciously, the accuracy post’s territory. It also explains why photo quality matters, a poor known-photo produces an off-centre encoding that real appearances hover at the edge of. And it explains speed, encoding is the heavy step, which is why it runs only at detected locations, and why the threading post’s frame management exists at all.

A few things people ask me about this

Why does my recognition run but match poorly? Check the BGR to RGB conversion first, it is the classic silent accuracy killer with OpenCV frames. Then check known-photo quality, one clear frontal face per person.

What does the tolerance value actually do? It is the maximum encoding distance that counts as a match, around 0.6 by default. Lower means stricter, fewer false accepts, more false rejects, tune it against your own people and lighting.

Next

One camera recognising faces is a demo, my requirement was two cameras live at once beside a responsive interface, which is a threading problem before it is a vision problem. That is the next post.

Leave a Reply

Your email address will not be published. Required fields are marked *