Building Kompro: An Open-Source, Self-Hosted Compliance Platform
Compliance management is one of those domains where organizations routinely hand over their most sensitive data to third-party SaaS vendors. Audit findings, risk assessments, security control statuses, evidence artifacts, these are the kinds of things that should probably stay within your own perimeter, not sit in someone else's database. Kompro exists because I wanted to demonstrate that a single developer could build a production-grade compliance platform that organizations actually own, run on their own infrastructure, and control end to end.
This is probably the journal notes of the architecture, the technical decisions behind it, the tradeoffs I made along the way, and honestly, a few things I'd reconsider if I started over.
The Shape of the Problem
A compliance platform has to do a lot of things at once, and none of them are optional. It needs to model multiple compliance frameworks, ISO 27001, SOC 2, GDPR, and whatever custom ones an organization has cooked up, without welding the data model to any single one of them. It needs to track policies, security controls, evidence, risk, incidents, and audits, all pointing at each other in sensible ways. It needs role-based access control that an organization can actually configure themselves, without opening a code editor. And it needs to store evidence files, some of which are genuinely sensitive, without those files leaking out to some external service the operator never agreed to.
The principle I settled on early, and kept coming back to whenever I got stuck, was this: the organization's compliance state is the source of truth, and frameworks are just different lenses over it. A control like "encrypt data at rest" might satisfy requirements in ISO 27001, SOC 2, and GDPR all at once. The data model needed to reflect that reality instead of duplicating the same control three times under three different framework labels.
The Shape of the Codebase
Kompro lives as a mono-repo with two halves: a backend running Express, Prisma, and PostgreSQL on port 5000, and a frontend running React, Vite, and Tailwind on port 5173. The root project just orchestrates the two of them running side by side. There's deliberately no shared root environment file. Each half loads its own configuration from its own directory, because both dotenv and Vite resolve environment variables relative to wherever they're actually run from, and trying to share one root file would have meant a bunch of path gymnastics for no real benefit.
The backend serves everything under a single API prefix. The frontend is a single-page app that proxies API calls through to the backend, both during development through Vite's dev server and in production through an nginx reverse proxy or Vercel rewrites. Keeping everything on the same origin like this matters more than it might sound, because authentication runs on httpOnly cookies, and keeping the UI and the API on the same domain sidesteps an entire category of cross-site cookie headaches before they even start.
Why Express, and Not Something Like NestJS or Fastify
Express is boring, in the best possible sense of that word. The dependency surface is small. The middleware model is something almost every backend developer already understands without a learning curve. For a project meant to be self-hosted and maintained by small teams, keeping the learning curve low matters more than any framework feature I could point to. The backend runs on just thirteen production dependencies, and that number is intentional, not an accident of not getting around to adding more.
Prisma won out over raw SQL or something like Knex because treating the schema as code gives you one single source of truth for migrations, type generation, and documentation all at once. The tradeoff is that Prisma's query engine is a Rust binary under the hood, which adds a bit to startup time and binary size. For a self-hosted server that starts once and then just runs, that cost basically disappears into the noise.
The Data Model
The schema ended up with 27 models, and a handful of structural decisions shaped almost everything else.
I left out enums entirely. Every status and type field is just a plain string with a default value. I considered Prisma enums seriously, but talked myself out of them for two reasons. Adding a new status value to a Prisma enum requires a migration, and in a self-hosted product, you want the operator to be able to extend the system on their own timeline rather than waiting on me to ship an upstream schema change. And since both the frontend and backend already treat statuses as display strings day to day, the type safety an enum would buy me in a JavaScript codebase is pretty marginal anyway. The cost is that nothing at the database level stops someone from inserting an invalid status. Application validation handles the normal path, but a typo in a seed script or a raw SQL query could still slip garbage in. For a single-tenant product where the operator is also the admin, I'm fine with that tradeoff.
The owner field on controls, policies, and a few other models is just a free-text string, not a foreign key into the user table. This came from a pretty practical observation: in real compliance work, the "owner" of a policy or control is often a department head or an external party who doesn't have, and may never have, an account in the system. Forcing a proper user relation there would have meant creating placeholder accounts or standing up a whole separate contacts table just to hold a name. Letting people type a name and move on was simpler. The cost is that you lose referential integrity, and you can't cleanly query "every control owned by this person" without doing string matching. The frontend softens that a bit with a searchable dropdown that fills the field in with a real user's display name when one exists, so the experience stays consistent even though the schema underneath isn't strictly coupled to it.
Frameworks map to controls through a chain: a framework has requirements, requirements connect to controls through a mapping record, and that mapping can carry its own notes. That means a single control can satisfy requirements from several frameworks at once, and framework readiness gets computed by walking that chain rather than being stored as some denormalized number that could drift out of sync. The readiness calculation checks, for every requirement, whether at least one mapped control is both implemented and has evidence attached, then rolls that up into a percentage per framework. The frontend turns that into a stacked bar chart and a prioritized list of gaps. The tradeoff is query complexity, since computing readiness means joining four tables together, and for an organization with thousands of requirements that could eventually become a real performance concern. I chose to compute it fresh every time rather than caching it in a materialized view, because the underlying data changes constantly during an assessment cycle, and a stale readiness number is a lot more dangerous than a slightly slow one.
One decision that tends to surprise people: policies aren't directly linked to frameworks at all. In practice, policies span multiple frameworks and organizational concerns simultaneously. A data retention policy might matter to GDPR, SOC 2, and a company's own internal governance all at the same time. Rather than build a many-to-many relationship between policies and frameworks that would complicate every single policy query, I left policies as standalone documents. The connection to frameworks flows instead through controls, since a control can cite a policy and that same control maps to framework requirements. Whether that was the right call really depends on the organization using it. Some compliance teams genuinely want to tag policies by framework directly, and the schema doesn't prevent adding that relation later. It's just not there today.
Authentication
Kompro authenticates with httpOnly session cookies rather than bearer tokens. The JWT gets set as a cookie marked httpOnly, secure in production, and same-site lax. It's never returned in a response body, and it never touches localStorage. This is a deliberate security choice: httpOnly cookies can't be read by JavaScript running on the page, which closes off an entire category of XSS-driven token theft. A bearer token sitting in localStorage is trivially exfiltrable by any script that manages to run in that page's context, and for a platform where a session grants access to audit findings and risk data, that risk isn't one I wanted to carry.
Session expiry runs on two layers. The JWT itself has a short lifetime, two hours by default. The underlying database session record has a much longer one, thirty days by default. Every authenticated request checks both: the JWT's own signature and expiry, and then the database session's revocation and expiry state. That combination gives you the performance benefit of JWTs, since most requests never need a database hit just to verify a signature, while still keeping real revocation available. When someone changes their password, or an admin deactivates an account, every active session for that account gets revoked instantly by flipping a flag on the database rows. The JWT itself might technically still be valid for up to two hours, but the database check catches it regardless. The cost is that every authenticated request does hit the database to check session status, which is fine for a self-hosted, single-tenant product with modest concurrency, but would need caching in Redis, or an accepted revocation delay, if this were a multi-tenant SaaS running at real scale.
Single sign-on through Google and Microsoft runs on OAuth 2.0's authorization code flow with PKCE for both providers. The PKCE verifier and the CSRF state get stored in a short-lived httpOnly cookie, ten minutes, rather than in server-side session storage. On the callback, the state gets verified, the code gets exchanged, the ID token gets decoded, and the user either matches an existing account by provider ID or email, or gets auto-provisioned with the standard member role. SSO accounts get a random bcrypt hash as their password, specifically so a password login attempt can never succeed against them by accident. The provider fields on the user model let accounts link automatically too, so if someone first signs up with email and password and later signs in with Google using that same email, the two accounts get tied together without any manual intervention.
Login attempts are rate-limited to five per fifteen minutes, keyed by email address rather than by IP. That choice matters more than it sounds: keying by IP would mean a shared corporate network, where hundreds of employees sit behind one NAT address, could get collectively locked out just because one person forgot their password. The rate limiter itself runs on an in-memory store, which means limits reset whenever the server restarts. For a single-instance deployment that's an acceptable gap. Scaling horizontally would need a Redis-backed store instead, and I deliberately didn't reach for that, because adding Redis as a dependency means asking every self-hosting operator, plenty of whom are running Kompro on a five-dollar VPS, to install, configure, and monitor one more piece of infrastructure.
Authorization
The role-based access control system is entirely data-driven. Permissions are rows in a table. Roles are rows in another table. The relationship between them is a straightforward many-to-many. On every request, the authorization middleware loads the user's role and that role's permissions straight from the database, which means permission changes take effect immediately, with no redeployment required. An admin can create a custom role, assign it a specific set of permissions, and any user with that role has those permissions on their very next request.
The seed script sets up 48 permissions spread across 12 domains, four CRUD operations per domain plus a few specials like collecting evidence or purging audit logs, and three default roles: an admin with everything, an auditor with read-only access, and a member limited to reading their own organization. The tradeoff here is a database query on every single protected request just to load the role's permission set. Prisma's connection pooling absorbs most of that cost, and at the scale this is built for, tens of concurrent users rather than thousands, the latency is basically imperceptible. If I ever needed to squeeze more performance out of it, caching role-permission mappings in memory with a short TTL, invalidated whenever a role changes, would be the obvious next step.
Storage
Evidence files land either on local disk or in S3-compatible object storage, decided automatically by whether the S3 bucket and region environment variables are both set. Every stored file's key gets prefixed with whichever driver stored it, so a file might be recorded as a local path or an S3 path, and that prefix lives right alongside the file record in the database. When the system needs to read or delete a file, it just parses that prefix to figure out which driver to use. That small detail means you can migrate from local disk to S3 later without rewriting a single existing database record. Old files keep their local prefix and keep being served from disk. New files get the S3 prefix and get served from the bucket instead. A migration script to move everything over would just need to upload each file and flip its prefix.
The local storage driver also checks that any resolved file path can't escape the upload directory, which closes off a class of path traversal attacks where a crafted filename might otherwise be used to read arbitrary files off the server. And the S3 client supports a path-style addressing mode for compatibility with MinIO and other S3-compatible stores, with credentials being optional so IAM role-based authentication works cleanly on AWS itself.
The Audit Trail
The audit log table is append-only, full stop. Every significant action, creating something, updating it, deleting it, logging in, changing a password, generating a reset link, gets recorded with the actor, the action, the entity type and ID, the client's IP address, and a before-and-after snapshot in JSON. There is no update function anywhere in the audit service, and that's completely deliberate. An audit trail that can be edited after the fact isn't really an audit trail anymore. There is a purge function for retention compliance, covering things like a GDPR right-to-erasure request or simple storage management, and it deletes entries older than a configurable number of days, 365 by default. Those before-and-after snapshots also let the frontend reconstruct a readable diff, rendered as a detail view so an administrator can see exactly what changed and when.
The Frontend
There's no state management library anywhere in the frontend. Just local component state and a single authentication context tracking the session. No Redux, no Zustand, no MobX, no React Query. Every page follows the same rhythm: fetch data on mount, call the API directly for any mutation, then refetch. There's no meaningful shared state across pages beyond who's currently logged in, so pulling in a state management library would have added real complexity, actions, reducers, selectors, cache invalidation, for a benefit that never really shows up in this particular architecture. The cost is that navigating between pages always triggers a fresh fetch, with no client-side cache sitting in between. For a compliance tool used by a small team on an internal network, that's a non-issue. For something with hundreds of concurrent users, you'd want something like React Query or SWR handling caching and background refetching instead.
All the shared UI components, buttons, badges, cards, spinners, page headers, empty states, modals, drawers, tables, and a handful of searchable select components, live together in a single file, a little over 400 lines long. The reasoning is that compliance tools have an extremely repetitive UI shape. Nearly every page is a table with some CRUD modals attached to it. Splitting each of those pieces out into its own file would have meant fifteen or more tiny files, each barely twenty or thirty lines long, adding directory-navigation overhead without actually making anything easier to read. Keeping it all in one module makes the entire design system visible at a glance and keeps everything visually consistent by construction. Vite's tree-shaking takes care of trimming unused exports, so the file staying large doesn't translate into a larger shipped bundle.
Visually, the app runs on a custom Tailwind theme built around two color scales, a cool near-black charcoal for the sidebar and primary actions, and a purple brand scale anchored around a specific shade as the primary color. The typeface is Inter with a serif fallback for headings, which gives the whole thing a slightly more editorial feel than the typical SaaS dashboard look. A handful of shared component classes for cards, inputs, labels, and badges live in the global stylesheet, mainly because certain utility combinations repeat across dozens of elements, and pulling them into named classes cuts down on a lot of HTML verbosity.
Deploying It
Kompro supports a few different deployment shapes. The simplest and most private is everything on a single Linux machine: PostgreSQL, the Express API, and nginx serving the built frontend, all on one host, with nginx reverse-proxying API calls to Express so everything stays on one domain and cookies behave correctly. A split setup works too, backend on Render, frontend on Vercel, database on something like Supabase or Neon, with Vercel's own rewrite rules routing API calls to the Render backend server-side so the browser only ever talks to the Vercel domain and the session cookie stays first-party. And beyond those two, any custom combination of hosting works, as long as the UI and API end up on the same registrable domain, or sit behind a proxy that makes them look that way, because the session cookie is hardcoded to the lax same-site setting.
That lax setting isn't configurable, and that was a deliberate choice too. Lax cookies get sent on same-site navigation and top-level GET requests but get blocked on cross-site fetch or XHR calls, which is exactly the right default for a self-hosted product: it blocks cross-site CSRF while still working normally for anyone using it the intended way. If someone deploys the frontend and backend on two completely unrelated domains without a proxy in between, the cookie simply won't get sent on API calls, and the fix is either to use a proxy, which is what I'd recommend, or to deliberately relax the cookie to none with the secure flag set, which the setup guide documents for anyone who really needs it. I chose to hardcode lax rather than exposing it as a setting, because a misconfigured none cookie set by someone who doesn't fully understand the security tradeoff is worse than a deployment that visibly doesn't work until the operator actually reads the documentation.
Migrations
Schema changes go through Prisma Migrate. A base migration, 749 lines of SQL, captures the initial schema, and everything after that flows through the normal "migrate dev" and "migrate deploy" commands in development and production respectively. I retired the quicker "db push" approach early on, since it doesn't produce migration files and can't be reliably reproduced across different environments. One genuinely painful lesson from this process: on Windows, PowerShell's redirection operator writes files as UTF-16LE by default, and Prisma's Rust-based migration engine can't parse that encoding, failing with error messages that give you almost no hint about what actually went wrong. If you're generating migration SQL on Windows, explicitly force UTF-8 encoding or convert the file before applying it, or you'll lose an afternoon to it the way I did.
What I'd Reconsider
The in-memory rate limiter resets on every server restart and doesn't work across multiple instances. For the single-server deployment this is built for, that's acceptable, but it's a real gap, and a Redis-backed limiter would close it properly. I skipped that mainly because Redis is one more thing to install, configure, and monitor, and the target operator here is genuinely someone running this on a cheap VPS, not an ops team.
There's no WebSocket layer anywhere. The frontend just fetches fresh data on page load, so if two administrators happen to be editing controls at the same time, neither one sees the other's changes until they refresh. For a small compliance team this rarely matters in practice, but it's a known limitation worth naming honestly.
The backend has solid integration test coverage, covering authentication flows, password resets, user management, and evidence CRUD. The frontend has none. React Testing Library and Playwright are the obvious next additions, and I prioritized shipping features over building that coverage, which is a tradeoff I'd genuinely revisit before calling this a 2.0.
There's no caching layer of any kind. Every authenticated request checks the database for session validity, every page load fetches everything fresh, and there's no Redis cache, no HTTP cache headers on API responses, and no client-side query cache. It works fine at the scale this was built for and would need real caching work before comfortably handling hundreds of concurrent users.
Password hashing currently runs bcrypt at 10 salt rounds. OWASP's minimum recommendation is 10, with 12 or higher suggested for genuinely sensitive applications. Bumping that to 12 would push hash time from around 100 milliseconds to roughly 400 milliseconds per login, which is a completely acceptable cost for a compliance platform where logins aren't a high-frequency event. It's on the roadmap.
The audit log table is append-only with a configurable retention window, but for organizations with heavy day-to-day activity, that table can grow quite large over time, and there's currently no partitioning or archival strategy beyond the periodic purge. Time-based partitioning in PostgreSQL would help here, but Prisma doesn't natively support partitioned tables, which makes this a harder problem than it initially looks.
And finally, the email templates currently reference a logo hosted on Cloudinary. If that service goes down, or the URL ever changes, every outgoing email quietly ends up with a broken image. A properly self-hosted alternative would inline the logo as a base64 data URI, or host it directly on the same domain as the application itself.
Closing
Kompro covers the full core compliance lifecycle: frameworks, controls, policies, evidence, assessments, risk, incidents, IT service management, and audit programs. It runs on a single PostgreSQL database, stores files on local disk or S3, and authenticates through secure session cookies rather than anything sitting in browser storage.
The project is open source under AGPL v3, and the source lives at github.com/KingDavidJnr/kompro, with self-hosting instructions, environment variable documentation, and deployment patterns all covered in the setup guide.
None of the decisions documented here are meant as universal best practices. They're tradeoffs made for one specific context: a single-tenant, self-hosted compliance tool built for small-to-medium organizations. Different constraints, multi-tenancy, enterprise scale, SaaS distribution, would have pushed me toward different answers almost every time. The point isn't that any of these choices are objectively correct. It's that every choice has a reason behind it, and knowing that reason is what actually makes a system maintainable years down the line.
Comments