7.5 KiB
Architecture
Marketing Parser is a single Spring Boot 3.5.5 service on Java 21, backed by MongoDB for persistence and MinIO for generated artefacts. It exposes a REST API consumed by the KonturAI frontend and runs a set of scheduled background jobs.
Functional areas
The service covers five loosely coupled concerns that share a database and a set of AI clients:
- RSS ingestion — collects business news into a
MarketItemcorpus. - Marketing analysis — turns that corpus plus user input into structured analysis.
- Report rendering — renders analysis as PDF, DOCX and Markdown with charts.
- Campaign execution — generates creatives and publishes them to social networks.
- Targeting and leads — ad targeting recommendations and Facebook lead capture.
Request flow
Frontend
│ Bearer JWT
▼
Controller ──validator──► DTO
│ │
│ ▼
│ Service layer
│ ╱ │ ╲
│ AI clients Repository MinIO
│ (OpenAI/Ollama/ │ (artefacts)
│ Vertex/Serper) ▼
│ MongoDB
▼
GlobalExceptionHandler ──► consistent error payload
Authentication is JWT bearer tokens issued by a separate auth service. This
service does not log users in; JwtService validates the signature against
security.jwt.secret-base64 and extracts the user id, email and roles from claims.
Controllers pull the caller identity via extractUserIdFromHeader(authHeader). See
api/authentication.md.
Layers
| Package | Responsibility |
|---|---|
controller |
HTTP endpoints, 15 controllers plus GlobalExceptionHandler |
validator |
Request-level validation beyond Bean Validation annotations |
dto |
Request/response payloads, including dto/targeting |
service |
All business logic — 52 classes |
repository |
Spring Data MongoDB interfaces |
model |
MongoDB documents |
config |
Beans, typed @ConfigurationProperties, async executors, CORS, Swagger |
exception |
Domain exceptions surfaced by GlobalExceptionHandler |
Endpoints
| Base path | Controller | Purpose |
|---|---|---|
/api/parser/health |
HealthCheckController |
Liveness |
/api/parser/items |
MarketItemController |
Access the ingested news corpus |
/api/parser/admin/parsers |
ParserAdminController |
Trigger and inspect parsers |
/api/parser/report |
ReportController |
Research report generation and history |
/api/marketing/analysis |
MarketingController |
Marketing analysis (v1/v2) |
/api/marketing/v3 |
MarketingAnalysisV3Controller |
Marketing analysis v3 |
/api/marketing/targeting |
TargetingCampaignController |
Campaign targeting |
/api/marketing |
PublicAssetController |
Public access to generated assets |
/api/targeting |
AiTargetingSystemController |
AI targeting recommendations |
/api/social-media/credentials |
SocialMediaCredentialsController |
Per-user network credentials |
/api/facebook/config |
FacebookConfigController |
Facebook app/page configuration |
/api/facebook/leads |
FacebookLeadsController |
Collected hot leads |
/api/facebook/webhook |
FacebookWebhookController |
Facebook webhook receiver |
/api/openai |
OpenAITestController |
Connectivity diagnostics |
Full request/response detail: api/README.md, or the live OpenAPI UI
at /swagger-ui.html.
RSS ingestion
ParserService is a small interface — getSourceName() and parseAndSaveRssFeed().
Each source implements it, and ParserManagerService acts as a facade: Spring injects
every ParserService bean and the manager indexes them by source name, so adding a
source requires no changes to the manager or the controller.
Five sources are implemented: Kursiv, Kapital, LSM, RBC and Vedomosti. Details and scheduling in rss-parsers.md.
Scheduled jobs
parser.scheduler.enabled is false by default — RSS polling does not run unless
explicitly enabled. posting.scheduler.enabled defaults to true.
| Job | Cron | Owner |
|---|---|---|
| Kursiv / Kapital ingest | every 5 min | KursivParserService, KapitalParserService |
| LSM / RBC / Vedomosti ingest | every 30 min | respective parser services |
| Scheduled post publishing | every minute | PostingSchedulerService |
| Facebook lead collection | every 15 min (configurable) | FacebookLeadCollectorService |
| Targeting campaign sync | every 6 hours | TargetingCampaignService |
| Campaign prediction refresh | daily 06:00 | CampaignPredictionService |
AI provider strategy
The service deliberately mixes providers by cost and capability:
- OpenAI (
gpt-4o,gpt-4o-mini) — structured JSON analysis and chart data, where reliable schema adherence matters. Wrapped in retry with exponential backoff and a concurrency cap of 3. - Ollama (self-hosted) — long-form narrative text, avoiding per-token cost on the largest outputs. Timeouts are correspondingly long (up to 5 hours).
- Google Vertex AI — Imagen 3 for images, Veo 3 for video, authenticated with a service-account key. See configuration.md.
- Serper — Google search results feeding research reports.
ClaudeApiService and DeepResearchService cover additional generation paths.
Persistence
MongoDB documents, one repository each:
| Document | Holds |
|---|---|
MarketItem |
Ingested news articles (the corpus) |
MarketingAnalysis, MarketingAnalysisV2Document, MarketingAnalysisV3Document |
Three analysis generations, kept side by side |
MarketingStrategy |
Generated promotion strategies |
TargetingCampaign, TargetingAudienceProfile, TargetingAdSet, TargetingAd, TargetingInsight |
Ad targeting model |
PostingTask |
Queued and published social posts |
SocialMediaCredentials |
Per-user network credentials, encrypted at rest |
FacebookLead |
Leads harvested from page comments |
ReportHistory |
Generated report metadata |
CampaignPrediction, ABTestConfig, BudgetConfig, PerformanceMetrics |
Campaign optimisation |
The three analysis document versions coexist because the API kept older revisions
working for the frontend; MarketingAnalysisV3Service is the current path.
External dependencies
| Dependency | Used for | Failure behaviour |
|---|---|---|
| MongoDB | All persistence | Fatal — service cannot operate |
| MinIO | Reports and generated images | Feature-level failure |
| OpenAI | Analysis, strategy, chart data | Retried, then surfaced as an error |
| Ollama | Long-form report text | Retried, then surfaced as an error |
| Vertex AI | Image/video generation | Logs error, returns null; rest of the service unaffected |
| Serper | Research search | Retried |
| Facebook Graph API | Posting and lead collection | Logged, retried on next schedule |
| SMTP | Emailing reports | Health indicator disabled; failures logged |
Cross-cutting configuration
- CORS —
CorsPropertiesbindscors.*, applied byWebCorsConfig. Origins default to the production frontend domains. - Async —
AsyncConfigandTargetingAsyncConfigprovide executors; long AI calls run off the request thread and parsers returnCompletableFuture. - Error handling —
GlobalExceptionHandlermaps domain exceptions to consistent payloads. - API docs —
SwaggerConfigdeclares thebearerAuthscheme so the Swagger UI can authorise with a JWT.