Here is how most content projects start. Someone decides the website needs cleaning up. Maybe the SEO team flagged a problem, maybe a rebrand is coming, maybe you are about to migrate. So a person opens HubSpot, starts clicking through pages, and begins fixing things as they go.
That is the wrong first move. Every time.
Before you change a single page, you need to know what you actually have. Not a rough idea. The real numbers. How many pages, how many blog posts, how many of them are missing a meta description, how many point to a canonical URL that no longer exists, how many are duplicates of each other. That is a content audit. And the reason projects run over budget and break things in production is almost always that nobody did one properly first.
What a content audit actually is
A content audit is a full inventory of your website content plus a health check on every record in it.
Most people think of it as a reading exercise. Open each page, judge whether the copy is still good, decide keep or kill. That is a content review, and it has its place. But it is not what saves you during a cleanup or a migration.
The audit that matters treats your content as data. You are not asking "is this page any good." You are asking "how many pages have no meta title," "which blog posts share the same slug pattern," "which redirects point to a 404," "which pages have no internal links pointing to them." Those are questions you answer with counts and filters, not with opinions.
If your audit output is a document full of notes, it is going to sit in a folder. If your audit output is a table you can sort and filter, it becomes a to-do list you can actually work through.
Why most HubSpot audits fail
There are two failure modes, and teams usually hit one of them.
The first is skipping it. The work feels urgent, the audit feels like overhead, so people jump straight to fixing. Then halfway through they discover the problem is three times bigger than they thought, or they realize they have been editing pages that were about to be deleted anyway. Wasted hours.
The second is the eyeball audit. Someone commits to being thorough, opens every page one at a time, and writes findings into a spreadsheet by hand. On a small site that is tedious but doable. On a portal with 500 pages, a few hundred blog posts, and a dozen HubDB tables, it falls apart. By the time you finish auditing, the early records are already out of date, and you have spent twenty hours producing a snapshot instead of fixing anything.
The problem in both cases is the same. HubSpot does not give you a single place to see all your content as structured data. Content lives across pages, blog posts, landing pages, authors, tags, redirects, and HubDB tables that were never designed for the load you are putting on them. Each of those lives in its own editor. So an audit means bouncing between screens, and nobody wants to do that by hand.
The reframe: audit content as structured data
Here is the shift that makes an audit useful. Stop thinking of your content as pages. Start thinking of it as rows in a table.
A blog post is a row. Its title, slug, meta title, meta description, author, tags, publish date, and canonical URL are columns. Once you see it that way, an audit stops being a reading task and becomes a filtering task. You are looking for empty cells, duplicate values, and patterns that should not exist.
This is the whole reason a spreadsheet view beats a page-by-page click-through. You can sort every blog post by meta description and instantly see the blank ones grouped together. You can filter for pages published before a certain date. You can spot ten different spellings of your product name in one scroll. None of that is possible when your content is locked behind individual page editors.
What to actually audit
Here is the checklist. Run every content type you have through it.
Content inventory and counts. Start with the raw numbers. How many pages, blog posts, landing pages, authors, tags, redirects, and HubDB rows exist. You cannot plan a project until you know the size of it, and the count almost always surprises people.
Missing metadata. Filter for records with no meta title or no meta description. These are the pages quietly costing you search visibility. Group them and you have your first fixable list.
Duplicate and near-duplicate content. Look for repeated titles, repeated slugs, and blog posts that say the same thing. Duplicates confuse search engines and inflate your page count for no reason.
Orphan pages. These are pages with no internal links pointing to them. Search engines struggle to find them, users never reach them, and they sit in your portal draining crawl budget. Most teams do not know they exist until an audit surfaces them.
Redirect health. Check that your redirects still point somewhere real. Redirect chains and redirects that end in a 404 are common and quietly damaging.
Canonical tags. Look for pages pointing to a canonical URL that is wrong, missing, or broken. This matters even more if you run multiple regions or languages, where a bad canonical can wreck your international SEO setup.
Inconsistent text and branding. Old product names, outdated company details, a phrase you no longer use. If it appears across dozens of pages, you want to know the exact count before you decide how to fix it.
Tag and author consistency. Messy tags and duplicate author records make your blog look sloppy and break your archive pages. Audit them like any other field.
How to run it without opening every page
You already know the manual version. Open a record, check the fields, note the problem, close it, repeat a few hundred times. That is the twenty-hour version.
The faster version is to pull your entire content set into one searchable table and audit by filtering instead of clicking. Fetch every content type into a single view, sort and filter across all of it, and let the empty cells and duplicates group themselves. What took a full day of clicking becomes an afternoon of filtering.
This is exactly what Smuves was built to do. It connects to your HubSpot CMS and pulls pages, blog posts, landing pages, authors, tags, redirects, and HubDB tables into one spreadsheet view. You can search across everything, filter for the gaps, and export the whole audit to CSV or Google Sheets so the rest of your team can review it. The audit stops being a chore and becomes a five minute pull.
What to do once the audit is done
The point of auditing as data is that your findings are already a work list. Each problem group maps to a fix.
Missing meta descriptions across 300 posts? That is a bulk metadata update, not 300 individual page edits. Old product name showing up everywhere? That is a find and replace across your entire site, not a manual hunt. And when you are making changes at that scale, doing it in a spreadsheet where you can review before you push is far safer than editing live pages one by one.
The audit tells you what is wrong. Bulk editing fixes it. Both work best when your content is data instead of a pile of pages.
The takeaway
Do not touch a single page until you know what you have. Skip the audit and you will discover the real scope halfway through, usually after you have already broken something. Do it by eyeballing every page and you will spend a day producing a snapshot that is stale before you finish.
Audit your content as structured data. Count it, filter it, group the gaps, and turn the findings into a fix list. That is the difference between a cleanup that finishes on time and one that quietly eats a month.
