How-To Guide

CV Parsing for Recruiting: How It Works and Where It Fails

CV parsing for recruiting extracts candidate data into an ATS. Learn where it fits in your workflow, which CV layouts break it, and what to check.

CV parsing for recruiting is the automated extraction of candidate data from a CV into structured fields inside an ATS, and its accuracy depends almost entirely on how the CV is formatted.

When a candidate's file arrives at your agency, software reads it and pulls out the name, contact details, work history, education, and skills, then stores them in the database fields your system can search. That is the first parse. The second parse happens inside your client's ATS when you submit the formatted CV for screening. A layout that confuses your parser is likely to confuse the client's parser too.

This guide covers where parsing sits in the source-to-submission workflow, which formatting choices cause data loss at each stage, and what to check before a CV leaves your hands. For a primer on how parsing works at a technical level, see CV parsing explained for recruiters.

Key takeaways

  • CV parsing for recruiting is the automated extraction of candidate data from a CV into structured fields inside an ATS, and its accuracy depends almost entirely on how the CV is formatted.
  • A CV goes through two separate parse events: your own ATS or CRM when the candidate first arrives, and the client's ATS when you submit the formatted CV for screening.
  • Tables, columns, graphics, and contact details placed in headers or footers are the formatting choices most likely to cause data loss at both parse points.
  • Each ATS has its own parser, and widely used systems like Workday and SuccessFactors each handle documents differently, so a format that passes one may fail another.
  • Run the text-copy test on every CV before submission: select all text, paste into a plain document, and read top to bottom. Scrambled or missing words means the client's ATS will have the same problem.
CV arrives at agency Your system reads it PARSE 1 Agency reformats into branded template Your control point Submitted to client Client ATS reads it PARSE 2
A submitted CV goes through two separate parse events. The reformatting step is where the agency has full control over what goes into the client's ATS.

Where parsing fits in your workflow

The typical agency workflow from candidate intake to client submission runs through five stages. A candidate submits a CV and your ATS or CRM reads it, extracts the data, and creates a candidate record. A recruiter then screens and qualifies the candidate using that record. The agency reformats the CV into a branded template. The formatted CV is sent to the client. The client's ATS reads the document before any hiring manager sees it.

Parsing is active at the first and last stages. The vast majority of large employers screen incoming applications through an applicant tracking system, so when you submit a CV to a client, you should assume software reads it first.

The reformatting step is where your agency has the most control. The document you reformat is the one that goes into the client's ATS. Getting the structure right at that point protects both parse events.

The double-parse problem

A recruitment agency faces the parsing problem twice, in a way that differs from a job seeker's experience. A job seeker worries about one ATS: the system on the other side of an application form. An agency faces the problem at intake and again at submission.

When a badly formatted CV arrives from a candidate, the first parse can fail silently. A phone number sits in a footer the parser skipped. A skills section appears as scrambled text because it was inside a table. The recruiter working from that record may not notice, and the wrong data travels forward into the reformatting step.

If the reformatted CV carries any of the original layout problems forward, the client's ATS hits the same issues. A 2021 report by Harvard Business School and Accenture found that 88% of employers surveyed believed their ATS was filtering out high-skilled candidates before a human reviewer ever saw them. A strong candidate can drop off a client's shortlist for a formatting reason, not a skills reason. When that happens, the agency that submitted the CV is associated with the loss.

Manual reformatting is slow, often taking anywhere from a few minutes to the better part of an hour per CV. Discovering a layout problem after reformatting means doing that work again. Catching it before intake means you only do the work once.

Formatting choices that cause data loss

These five elements appear regularly in candidate CVs and cause data loss at one or both parse points. Identifying and fixing them before reformatting prevents the problem from reaching the client.

Tables and multi-column layouts

ATS parsers read linearly, left to right and top to bottom. When a CV uses a table or two columns, the parser slices horizontally across the full page rather than reading each section in sequence. A design-heavy CV built with tables and columns can lose entire sections this way: skills, contact information, the About section, and portfolio links placed in a second column are read out of order or dropped entirely.

Graphics, logos, and skill bar charts

A parser reads text, not images. Any information encoded as a graphic, logo, icon, or skill progress bar does not appear in the parsed output. A candidate whose skills are displayed as a visual chart will have no skills in the candidate record after parsing. This is true whether the graphic is decorative or carries real data.

Contact details in headers or footers

Many parsers skip or drop document headers and footers entirely. Contact details stored in a header or footer are frequently missed, so a candidate can appear unreachable in the system even though their number was clearly visible on the page.

Non-standard section headings

Parsers look for expected labels such as Work Experience, Education, and Skills to sort content into the right fields. A heading like "My Story" or "Where I've Been" leaves the parser guessing. It may classify the content in the wrong field or skip the section entirely. Standard labels keep the sorting step reliable.

Scanned or image-based PDFs

A scanned CV is a photograph, not a text document. Without OCR, the parser has nothing to read and every field comes back empty. Image-only PDFs and CVs exported as flat images carry the same problem. A text-based PDF (where the text is selectable) or an editable .docx file gives the parser real text to work with.

How to check a CV before you submit it

Run through these five checks for every CV, in order. The first three happen before intake or before reformatting. The last two happen after your system has processed the file.

Step 1: Check the file type before anything else

If the file is a scanned PDF or an image-based document, the text layer is missing and OCR must run before parsing can work. A text-based PDF (where the text is selectable when you click it) and a .docx file both give the parser real text to read. Flag or convert image-only files before they reach your CRM.

Step 2: Run the text-copy test

Open the CV, select all the text with Ctrl A or Command A, copy it, and paste it into a plain text document. Read the result top to bottom. If the words are scrambled, out of order, or entire sections are missing, the parser will have the same problem. This test catches multi-column and table issues immediately, without any specialist tool.

Step 3: Confirm contact details are in the body

Check that the candidate's name, phone number, and email address appear in the main body of the document as plain text. If they sit inside a floating header, a text box, or a footer section, move them into the body before reformatting. A contact detail the parser cannot find is a direct placement risk.

Step 4: Check that section headings are standard

Scan the CV for section labels. The headings should be immediately recognisable: Work Experience or Professional Experience, Education, Skills, Certifications. Rename any creative or personalised heading before you reformat. Your branded template should already use standard labels, so this check mainly targets the incoming source document.

Step 5: Review the parsed output field by field

After your system ingests the CV and creates a candidate record, open the record and check each key field. Look at the name, email address, phone number, current job title, current employer, skills, and qualifications. Anything wrong or missing must be corrected before the record or the CV moves to the next step. This review takes two minutes and catches the errors that would otherwise travel forward.

If you reformat CVs at volume, RefineCV handles the file-type check, OCR for scanned documents, and header and footer cleanup automatically, then rebuilds each CV into a clean single-column branded template. You review the parsed fields before export. Try it free with 10 CVs, no card. See transparent pricing.

What the parser most often misses

After your system ingests a CV, open the candidate record and check each of these fields. These are the ones most likely to be wrong, empty, or placed in the wrong category after parsing.

Check these fields after every intake

  • Name: confirm the spelling matches the CV exactly
  • Email address and phone number: confirm both are present and correct
  • Current job title: confirm it matches the most recent role on the CV
  • Current employer: confirm the company name was captured correctly
  • Skills: confirm the section was not dropped due to a table or graphic format in the source CV
  • Qualifications and certifications: confirm these were not lost when placed in a sidebar or image in the source
  • Employment dates for each role: confirm they parsed as dates, not as stray text or numbers
  • About or profile summary: confirm it was not dropped due to a two-column or graphically heavy source layout

Common mistakes to avoid

Forwarding a CV without running the text-copy test

If you have not checked the text layer, you do not know what the client's ATS will receive. The test takes under a minute and catches the most common layout failures before they cost a placement.

Keeping a two-column layout in the reformatted CV

Reformatting is the point where you have full control over the structure. If the candidate's source CV was two-column, convert it to a single column in your branded template. Carrying the layout problem forward means the second parse fails for exactly the same reason the first one did.

Leaving contact details in the document header

A header looks like the top of the page, but it is a separate document layer that many parsers skip. Move the name, phone number, and email into the body of the document. Your branded template should already do this. Check that the source document also has these details in the body before you transfer data.

Using icons or graphic elements for contact details or skills

Phone icons, email symbols, and skill progress bars look polished in a template but break parsing for the same reason a full graphic does: there is no text for the parser to read. Use plain text labels and standard bullet points instead. Every piece of important information should exist as real, selectable text.

Assuming the same format is safe for every client

Widely used systems like Workday and SuccessFactors each handle documents differently, so a format that works for one client's system may not work for another's. The safest format for any client is single-column, text-based, with standard headings. That format gives every parser the best possible chance.

Frequently asked questions

What is CV parsing for recruiting?

CV parsing for recruiting is the automated process of reading a candidate's CV and extracting the relevant data into structured fields inside an ATS or CRM. The software identifies the name, contact details, work history, education, and skills, then loads them into database fields the recruiter can search and filter. The output is a structured format, often XML or JSON, which then imports into an ATS or HRIS. For recruitment agencies, parsing happens twice: when the candidate's CV first arrives at the agency, and again when the formatted CV is submitted to the client.

Why do tables and columns break ATS parsing?

ATS parsers read documents as a continuous linear stream, moving left to right and top to bottom. When a CV uses a table or two columns, the parser slices horizontally across the full width of the page instead of reading each section independently. This produces scrambled output: the job title from one column ends up mixed with the skill from the other, and dates appear beside random words. A design-heavy CV built with tables and columns can lose entire sections this way, with skills, contact details, the About section, and portfolio links in a second column read out of order or dropped.

Which CV formatting choices cause the most data loss?

The five most common causes are tables and multi-column layouts, graphics or visual elements where text should be, contact details placed in document headers or footers, non-standard section headings, and scanned or image-based PDFs with no text layer. Contact information stored in a header or footer is frequently missed, and tables, columns, and graphics are among the most common formatting problems that disrupt parsing.

How do I check whether a CV has been parsed correctly?

Run two checks. Before intake, do the text-copy test: select all the text in the file, paste it into a plain document, and read it top to bottom. If the words are scrambled or sections are missing, fix the layout before the file is ingested. After intake, open the candidate record your system created and check each key field: name, email, phone, current job title, current employer, skills, and qualifications. Correct anything that is wrong or missing before the record or the CV moves forward.

Does the same CV format work for every ATS?

No. Each ATS has its own parser and handles documents differently. Widely used systems like Workday and SuccessFactors each handle documents in their own way. A format that passes cleanly in one may fail in the other. Because you cannot always know which system a client uses, the safest approach is single-column, real selectable text, standard section headings, contact details in the body, and a text-based PDF or .docx file. That combination gives every major parser the best possible chance of extracting the data correctly.

The bottom line

CV parsing for recruiting runs at two points in every agency's workflow: when the candidate's CV arrives and your system reads it, and when the formatted CV reaches the client's ATS. Both depend on the same things: real selectable text instead of graphics, a single column instead of a table layout, standard section headings, and contact details in the body of the document rather than in a header or footer. The reformatting step is where you have full control over what goes out. Run the text-copy test before intake, review the parsed fields before anything moves forward, and reformat into a clean single-column output before you submit. That sequence protects the candidate's data at both ends.

Related reading: how to make a candidate CV ATS-friendly and the difference between an ATS and a recruitment CRM.

Stop losing candidate data to broken CV layouts

RefineCV parses CVs with OCR support, rebuilds them into clean single-column branded templates, and exports ATS-readable PDFs or Word files. Start free with 10 CVs, no credit card. Then $0.40 per CV, or $50 per month for 200.

Start Free, 10 CVs

Sources

The RefineCV Team

Written by the team building RefineCV, CV formatting software for recruitment agencies.

Format your next CV in 10 seconds

Try RefineCV with 10 free CVs. No credit card. Then $0.40 per CV, or $50/month for 200 on Pro.

Start Free, 10 CVs

No credit card required