Re: [DISCUSS] OpenCrawling Proposal

Piergiorgio Lucidi <[email protected]>
Newsgroups gmane.comp.apache.incubator.general
Message-ID <CAEO2op8Z-=wM_N0c3NGJb24t1dZ7kq0WFP71izz8BUYLt_+Zjw@mail.gmail.com>
Il giorno mer 12 ago 2026 alle ore 15:34 tison <[email protected]> ha
scritto:

>
> https://cwiki.apache.org/confluence/spaces/INCUBATOR/pages/446071456/OpenCrawling+Proposal
>
> This page is ready. May you check if you have permission to edit this page?
> Note that you may log in with your Apache ID.
>

Unfortunately I can't edit the page with my Apache account.


>
>
> Best,
> tison.
>
>
> Piergiorgio Lucidi <[email protected]> 于2026年8月12日周三 21:31写道:
>
> > Hi tison,
> >
> > Thank you so much for publishing the proposal and for trying to solve
> this
> > but I'm continuing to have the same issue.
> > I tried to create a blank page giving Proposals as the main page but I
> see
> > the same permission error.
> >
> > Cheers,
> > PG
> >
> >
> >
> > Il giorno mer 12 ago 2026 alle ore 03:00 tison <[email protected]> ha
> > scritto:
> >
> > > Hi Piergiorgio,
> > >
> > > I've invited you to the page. But I'm unfamiliar with Confluence, so it
> > may
> > > not be what you're looking for.
> > >
> > > Alternatively, you may find
> > > https://selfserve.apache.org/confluence-account.html helps.
> > >
> > > Anyway, I created the proposal page [1] with the wiki page content you
> > > shared.
> > >
> > > [1]
> > >
> > >
> >
> https://cwiki.apache.org/confluence/spaces/INCUBATOR/pages/446071456/OpenCrawling+Proposal
> > >
> > > Best,
> > > tison.
> > >
> > >
> > > Piergiorgio Lucidi <[email protected]> 于2026年8月12日周三 03:36写道:
> > >
> > > > Hi PJ,
> > > >
> > > > Thank you for the feedback and for highlighting Apache StormCrawler.
> > > >
> > > > I completely agree that exploring collaboration between the two
> > projects
> > > is
> > > > a great idea.
> > > >
> > > > While StormCrawler is an incredibly powerful tool for large scale web
> > > > crawling and processing, OpenCrawling was built specifically to
> tackle
> > > the
> > > > enterprise content and process automation ecosystem.
> > > >
> > > > Our core focus is bridging platforms like Alfresco, Flowable, and
> > Camunda
> > > > directly into modern LLM and RAG architectures, heavily utilizing
> > Spring
> > > > Boot and Spring AI.
> > > >
> > > > Despite the distinct use cases, web data versus enterprise
> > repositories,
> > > > there is absolutely a shared interest in robust data ingestion,
> > document
> > > > parsing and vector database integration strategies.
> > > >
> > > > OpenCrawling can become the perfect home for any crawling strategy.
> > > >
> > > > I would be thrilled to connect with the StormCrawler community to see
> > how
> > > > our projects might complement each other and share best practices
> > moving
> > > > forward.
> > > >
> > > > We could propose to implement a brand new OpenCrawling Storm Bolt.
> > > >
> > > > In Apache Storm topology, data flows from *Spouts* (URL queues) to
> > > *Bolts*
> > > > (fetchers, parsers and indexers). The most seamless integration is to
> > > build
> > > > an opencrawling-storm-bolt.
> > > >
> > > > StormCrawler handles the heavy lifting of recursive web crawling,
> > > > politeness and HTML parsing.
> > > >
> > > > Instead of using StormCrawler's native OpenSearch/Elasticsearch
> indexer
> > > > bolt, the topology passes the parsed document to the new custom
> > > > OpenCrawling Storm Bolt.
> > > >
> > > > The Bolt acts as an OpenCrawling Repository Connector. It takes the
> raw
> > > > HTML, wraps it in the Open Ingestion Standard (OIS) format, attaches
> > any
> > > > relevant baseline metadata and pushes it through OpenCrawling’s
> secure
> > > > pipeline (Java 25 / Spring Boot 4) for chunking, embedding, and
> Vector
> > DB
> > > > ingestion.
> > > >
> > > > Cheers,
> > > >
> > > > PG
> > > >
> > > > Il Mar 11 Ago 2026, 18:44 PJ Fanning <[email protected]> ha
> > scritto:
> > > >
> > > > > This doesn't block OpenCrawling joining as an ASF Incubator podling
> > > > > but I just want to highlight that there is already Apache
> > > > > StormCrawler.
> > > > >
> > > > > https://stormcrawler.apache.org/
> > > > >
> > > > > It would be great if these projects could collaborate in areas of
> > > > > shared interest.
> > > > >
> > > > > On Tue, 11 Aug 2026 at 15:28, Piergiorgio Lucidi <
> > > [email protected]
> > > > >
> > > > > wrote:
> > > > > >
> > > > > > Hi everyone,
> > > > > >
> > > > > > I would like to propose OpenCrawling as a new project for
> > incubation
> > > > > within
> > > > > > the Apache Software Foundation. Currently the project is hosted
> on
> > > > GitHub
> > > > > > [1].
> > > > > >
> > > > > > The official proposal is currently available in the OpenCrawling
> > Wiki
> > > > in
> > > > > > markdown format [2].
> > > > > >
> > > > > > I tried to share the proposal in our Confluence but it seems
> that I
> > > > don't
> > > > > > have permission to create the new page under the Proposals page.
> > > Anyway
> > > > > if
> > > > > > someone can guide me on resolving this issue it would be great!
> > > > > >
> > > > > > Once space permissions are granted on cwiki.apache.org, I will
> > also
> > > > > mirror
> > > > > > the proposal on the Incubator CWIKI proposals page.
> > > > > > Below you also find the same proposal ready to be copy-pasted
> into
> > > our
> > > > > > Confluence.
> > > > > >
> > > > > > We welcome feedback, questions and discussion from the Incubator
> > > > > community!
> > > > > >
> > > > > > Best regards,
> > > > > > Piergiorgio
> > > > > > On behalf of the OpenCrawling Core Team
> > > > > >
> > > > > > [1] - https://github.com/opencrawling/opencrawling
> > > > > > [2] -
> > > > > >
> > > > >
> > > >
> > >
> >
> https://github.com/opencrawling/opencrawling/wiki/Apache-Incubator-Proposal
> > > > > >
> > > > > > --------------------------------------------------------------
> > > > > >
> > > > > > h1. Apache Incubator Proposal: OpenCrawling
> > > > > >
> > > > > > h2. Abstract
> > > > > >
> > > > > > *OpenCrawling* is an open-source, enterprise-grade,
> > high-performance
> > > > data
> > > > > > crawling, content ingestion, and security-aware vector search
> > > platform.
> > > > > > Built on modern Java 25 (leveraging Virtual Threads and
> Structured
> > > > > > Concurrency), Spring Boot 4, and Spring AI, OpenCrawling serves
> as
> > > the
> > > > > > reference implementation of the *Open Ingestion Standard (OIS)*
> and
> > > > > > provides a secure *Model Context Protocol (MCP)* server
> interface.
> > It
> > > > > > orchestrates scalable data flows from heterogeneous enterprise
> > > > > repositories
> > > > > > (e.g., SharePoint, S3, CMIS/Alfresco, BPMN engines like Camunda
> and
> > > > > > Flowable, Apache Iceberg, Apache Ozone) to downstream vector
> > > databases
> > > > > > (e.g., Milvus, Qdrant, OpenSearch, Vespa, pgvector) with
> > source-level
> > > > > > Access Control List (ACL) security enforcement.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Proposal
> > > > > >
> > > > > > The OpenCrawling community proposes to incubate *OpenCrawling*
> as a
> > > new
> > > > > > project within the Apache Software Foundation (ASF). OpenCrawling
> > > > > provides
> > > > > > a decoupled, vendor-neutral enterprise data integration framework
> > > that
> > > > > > bridges the gap between traditional enterprise content management
> > > (ECM)
> > > > > > repositories and modern Large Language Model (LLM) /
> > > > Retrieval-Augmented
> > > > > > Generation (RAG) architectures.
> > > > > >
> > > > > > The project encompasses:
> > > > > > # *Core Ingestion Runtime ({{oc-core}}, {{oc-runtime}})*: A
> > > > distributed,
> > > > > > asynchronous engine built with virtual threads, claim-check
> > metadata
> > > > > > patterns, and Apache Kafka event streams.
> > > > > > # *Repository & Vector Connectors*: Standardized connectors for
> > > > scanning
> > > > > > source systems and indexing vector embeddings into major vector
> > > stores.
> > > > > > # *Open Ingestion Standard (OIS)*: The formal JSON/YAML schema
> > > > > > specifications defining unified document payloads, ACL security
> > SIDs,
> > > > and
> > > > > > crawler job configurations.
> > > > > > # *Secure Model Context Protocol (MCP) Server*: A Zero-Trust
> > context
> > > > > > retrieval server enforcing document-level permissions (Active
> > > Directory
> > > > > > SIDs, LDAP groups, user principals) at query time.
> > > > > > # *Observability & Developer Tooling*: AI-Powered Observability
> > > (AIOps)
> > > > > > over OpenTelemetry traces, Auto-Narrativization Copilot, Java
> > Client
> > > > SDK,
> > > > > > and Maven Archetypes for custom connector development.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Background
> > > > > >
> > > > > > In enterprise AI and RAG architectures, LLM agents require
> seamless
> > > > > access
> > > > > > to unstructured content stored across legacy and cloud
> > repositories.
> > > > > > However, traditional ingestion pipelines often strip out or
> ignore
> > > > > > source-level security metadata (ACLs), leading to context leakage
> > > where
> > > > > an
> > > > > > AI model synthesizes responses using confidential documents that
> > the
> > > > > > requesting user does not have permissions to view.
> > > > > >
> > > > > > Furthermore, legacy crawling tools (such as Apache ManifoldCF or
> > > Apache
> > > > > > Nutch) were architected over a decade ago prior to the emergence
> of
> > > > > vector
> > > > > > databases, LLMs, Model Context Protocol (MCP), and modern Java
> > > features
> > > > > > like Virtual Threads (JEP 444) and Structured Concurrency.
> > > > > >
> > > > > > OpenCrawling was created to address this modern ingestion crisis
> by
> > > > > > providing a native Java 25/Spring AI implementation engineered
> > > > > specifically
> > > > > > for LLM search scenarios, zero-trust context retrieval, and
> > > > > high-throughput
> > > > > > asynchronous processing.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Rationale
> > > > > >
> > > > > > The Apache Software Foundation is the natural home for
> > OpenCrawling.
> > > > ASF
> > > > > > has long been the center of innovation for enterprise search and
> > big
> > > > data
> > > > > > infrastructure, hosting cornerstone projects such as Apache
> Lucene,
> > > > > Apache
> > > > > > Solr, Apache Tika, Apache Kafka, Apache Iceberg, Apache Ozone,
> > Apache
> > > > > > ManifoldCF, and Apache Nutch.
> > > > > >
> > > > > > Bringing OpenCrawling to the ASF offers multiple mutual benefits:
> > > > > > * *Ecosystem Integration*: OpenCrawling directly integrates with
> > and
> > > > > builds
> > > > > > upon existing Apache projects, including *Apache Tika* (text
> > > > extraction),
> > > > > > *Apache Kafka* (event-driven pipeline), *Apache Iceberg*
> (lakehouse
> > > > > > connector), *Apache Ozone* (claim-check object storage), and
> > *Apache
> > > > > Maven*
> > > > > > (connector archetype distribution).
> > > > > > * *Vendor-Neutral Governance*: Neutral governance under the
> Apache
> > > Way
> > > > is
> > > > > > vital to establishing OpenCrawling and OIS as industry-wide,
> > > > > > vendor-agnostic ingestion standards.
> > > > > > * *Community Sustainability*: Operating as an Apache project will
> > > > > attract a
> > > > > > broader community of enterprise adopters, cloud providers, AI
> > > framework
> > > > > > developers, and search engine vendors.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Initial Goals
> > > > > >
> > > > > > During incubation, the OpenCrawling project will focus on the
> > > following
> > > > > > milestones:
> > > > > >
> > > > > > # *ASF Migration & Infrastructure*:
> > > > > > ** Transfer codebases ({{opencrawling}},
> > {{open-ingestion-standard}},
> > > > {{
> > > > > > opencrawling.github.io}}) to Apache infrastructure ({{
> > > > > > github.com/apache/incubator-opencrawling}}
> > > > > <http://github.com/apache/incubator-opencrawling%7D%7D>).
> > > > > > ** Rebrand build artifacts to {{org.apache.opencrawling}}.
> > > > > > ** Setup ASF-compliant CI/CD pipelines using GitHub Actions.
> > > > > > # *Community & Governance*:
> > > > > > ** Adopt the Apache Way for all decisions, roadmap discussions,
> and
> > > > > release
> > > > > > voting.
> > > > > > ** Expand the contributor base across independent developers,
> > > > enterprise
> > > > > > search users, and corporate contributors.
> > > > > > # *Ecosystem & Connector Expansion*:
> > > > > > ** Release additional output connectors (Elasticsearch, Apache
> > Solr,
> > > > > > RESTHeart).
> > > > > > ** Add native integration for fine-grained authorization
> frameworks
> > > > > (e.g.,
> > > > > > OpenFGA).
> > > > > > ** Enhance gRPC support for high-efficiency inter-microservice
> > > > > > communication.
> > > > > > ** Standardize OIS specification drafts under ASF governance.
> > > > > > # *Compliance & Licensing*:
> > > > > > ** Complete IP clearance and execute software grant agreements.
> > > > > > ** Ensure all third-party dependencies strictly conform to Apache
> > > > License
> > > > > > Category A policies.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Current Status
> > > > > >
> > > > > > h3. Meritocracy
> > > > > > The OpenCrawling project was established with meritocratic
> > principles
> > > > > from
> > > > > > day one. Design decisions, architecture changes, issue tracking,
> > and
> > > > > > roadmap discussions take place openly on GitHub through RFCs,
> Pull
> > > > > > Requests, and public wiki pages.
> > > > > >
> > > > > > h3. Community
> > > > > > The OpenCrawling community includes developers and architects
> from
> > > > > > enterprise search, ECM, and AI background. Community channels
> > include
> > > > > > GitHub Discussions, Slack, and social media announcements. The
> > > project
> > > > > > actively encourages external contributions via Maven archetypes
> and
> > > > > modular
> > > > > > connector development.
> > > > > >
> > > > > > h3. Core Developers
> > > > > > The initial core developers are experienced software architects
> and
> > > > > > open-source veterans with extensive experience in enterprise
> > search,
> > > > > > content management, and ASF governance:
> > > > > >
> > > > > > * *Piergiorgio Lucidi* ({{[email protected]}}) – Founder,
> > Lead
> > > > > > Architect. ASF Member and PMC Member/Committer on multiple Apache
> > > > > projects
> > > > > > (including Apache ManifoldCF and Apache Chemistry).
> > > > > > * *Michael Cizmar* ({{[email protected]}}) – Lead
> > Architect
> > > &
> > > > > > Developer. Specialist in enterprise search and cloud
> > infrastructure.
> > > > > > * *Luis Cabaceira* ({{[email protected]}}) – Lead
> > Architect &
> > > > > > Developer. Specialist in document processing and AI integration.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Known Risks
> > > > > >
> > > > > > h3. Orphaned Products
> > > > > > The risk of OpenCrawling becoming orphaned is low. The project
> > solves
> > > > an
> > > > > > active, urgent security and performance problem in enterprise AI
> > > > adoption
> > > > > > (RAG ACL context leakage). The core maintainers are committed to
> > its
> > > > > > long-term evolution and actively use it in production
> environments.
> > > > > >
> > > > > > h3. Inexperience with Open Source
> > > > > > The project leadership has deep experience with open-source
> > > > communities.
> > > > > > Piergiorgio Lucidi is an active ASF Member and PMC member with
> > over a
> > > > > > decade of experience guiding projects through the Apache Way.
> > > > > >
> > > > > > h3. Homogenous Developers
> > > > > > The initial committers come from diverse geographical locations
> > > (Italy,
> > > > > > United States, Portugal) and distinct
> organizations/consultancies.
> > > > > > Incubating at Apache will further diversify the developer base by
> > > > > > encouraging contributions from enterprise organizations and
> search
> > > > > vendors.
> > > > > >
> > > > > > h3. Reliance on Third-Party Products
> > > > > > OpenCrawling is designed to be vendor-neutral. Core dependencies
> > are
> > > > > > open-source libraries under permissive licenses (Apache 2.0, MIT,
> > > BSD):
> > > > > > * Spring Boot & Spring AI (Apache 2.0)
> > > > > > * Apache Tika (Apache 2.0)
> > > > > > * Apache Kafka (Apache 2.0)
> > > > > > * PostgreSQL / pgvector (PostgreSQL License / MIT)
> > > > > > * Docker & OpenTelemetry (Apache 2.0)
> > > > > >
> > > > > > There are no GPL/AGPL dependencies in the runtime core.
> > > > > >
> > > > > > h3. Relationship with Sponsored Products / Brand
> > > > > > OpenCrawling is an independent project. The name "OpenCrawling"
> has
> > > > been
> > > > > > used for the open-source codebase. The trademark will be
> > transferred
> > > to
> > > > > the
> > > > > > Apache Software Foundation upon incubation acceptance.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Documentation & Existing Artifacts
> > > > > >
> > > > > > * *GitHub Organization*: [https://github.com/opencrawling]
> > > > > > * *Main Code Base*: {{opencrawling/opencrawling}}
> > > > > > * *Specification Repo*: {{opencrawling/open-ingestion-standard}}
> > > > > > * *Documentation & Wiki*: [
> > > > > https://github.com/opencrawling/opencrawling/wiki
> > > > > > ]
> > > > > > * *Java Client SDK*: {{oc-java-client-sdk}} ([Sonatype Central -
> > > > > > org.opencrawling:oc-java-client-sdk|
> > > > > >
> > > > >
> > > >
> > >
> >
> https://central.sonatype.com/artifact/org.opencrawling/oc-java-client-sdk
> > > > > ])
> > > > > > * *Maven Archetypes*: [Sonatype Central -
> > > org.opencrawling.archetypes|
> > > > > >
> > > > >
> > > >
> > >
> >
> https://central.sonatype.com/artifact/org.opencrawling.archetypes/opencrawling-connector-archetypes
> > > > > > ]
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Initial Source & Intellectual Property Submission
> > > > > >
> > > > > > h3. Initial Source Code
> > > > > > The initial codebase to be granted to the ASF resides in the
> > > following
> > > > > > GitHub repositories:
> > > > > > * {{opencrawling/opencrawling}} (Core engine, microservices, UI,
> > > > > > connectors, MCP server)
> > > > > > * {{opencrawling/open-ingestion-standard}} (JSON schemas,
> > whitepaper,
> > > > > > specifications)
> > > > > > * {{opencrawling/opencrawling.github.io}} (Project web site and
> > > > > > documentation source)
> > > > > >
> > > > > > All source code is currently licensed under the *Apache License,
> > > > Version
> > > > > > 2.0*.
> > > > > >
> > > > > > h3. Software Grant / ICLA / CCLA
> > > > > > All core contributors will submit Individual Contributor License
> > > > > Agreements
> > > > > > (ICLAs) and corporate software grants will be executed upon
> > > acceptance
> > > > > into
> > > > > > the Incubator.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. External Dependencies
> > > > > >
> > > > > > All major external dependencies of OpenCrawling use
> > Apache-compatible
> > > > > > licenses (Category A):
> > > > > >
> > > > > > || Dependency || License ||
> > > > > > | *Java Development Kit (JDK 25)* | GPLv2 + Classpath Exception |
> > > > > > | *Spring Boot / Spring AI* | Apache License 2.0 |
> > > > > > | *Apache Tika* | Apache License 2.0 |
> > > > > > | *Apache Kafka Clients* | Apache License 2.0 |
> > > > > > | *Apache Iceberg SDK* | Apache License 2.0 |
> > > > > > | *Apache Ozone Client* | Apache License 2.0 |
> > > > > > | *Jackson / Slf4j / Logback* | Apache 2.0 / MIT / EPL 1.0 |
> > > > > > | *Milvus / Qdrant Java SDKs* | Apache License 2.0 |
> > > > > > | *OpenTelemetry Java SDK* | Apache License 2.0 |
> > > > > > | *React / Vite / Tailwind (Admin UI)* | MIT |
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Cryptography
> > > > > >
> > > > > > OpenCrawling uses standard TLS/HTTPS protocols and hashing
> routines
> > > > > > provided by the standard Java Virtual Machine (JDK) and Spring
> > > Security
> > > > > > framework for secure transport. It does not include custom
> > > > cryptographic
> > > > > > algorithms or controlled export software.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Required Resources
> > > > > >
> > > > > > h3. Mailing Lists
> > > > > > * {{[email protected]}}
> > > > > > * {{[email protected]}}
> > > > > > * {{[email protected]}} (PPMC)
> > > > > >
> > > > > > h3. Git Repositories
> > > > > > * {{https://github.com/apache/incubator-opencrawling}}
> > > > > > * {{https://github.com/apache/incubator-opencrawling-site}}
> > > > > >
> > > > > > h3. Issue Tracking
> > > > > > * GitHub Issues on {{apache/incubator-opencrawling}} (or ASF Jira
> > > > project
> > > > > > {{OPENCRAWLING}})
> > > > > >
> > > > > > h3. CI/CD Infrastructure
> > > > > > * GitHub Actions workflows for automated build, test, multi-arch
> > > Docker
> > > > > > image generation, and Sonar/Scorecard quality checks.
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Initial Committers & PPMC Members
> > > > > >
> > > > > > * *Piergiorgio Lucidi* ({{[email protected]}}) – Initial
> > > > Committer
> > > > > &
> > > > > > PPMC
> > > > > > * *Michael Cizmar* ({{[email protected]}}) – Initial
> > > > Committer &
> > > > > > PPMC
> > > > > > * *Luis Cabaceira* ({{[email protected]}}) – Initial
> > > Committer
> > > > &
> > > > > PPMC
> > > > > >
> > > > > > _(Note: Additional mentors and committers will be welcomed during
> > the
> > > > > > discussion period on {{[email protected]}}.)_
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Champions & Mentors
> > > > > >
> > > > > > * *Champion*: Piergiorgio Lucidi ({{[email protected]}}) –
> > ASF
> > > > > Member
> > > > > > * *Mentors*:
> > > > > > ** _(TBD - Interested ASF Members/Incubator PMC members invited
> to
> > > step
> > > > > > forward during proposal discussion)_
> > > > > >
> > > > > > ----
> > > > > >
> > > > > > h2. Sponsoring Entity
> > > > > >
> > > > > > The *Apache Incubator PMC* is requested to be the sponsoring
> entity
> > > for
> > > > > > this project.
> > > > >
> > > > >
> ---------------------------------------------------------------------
> > > > > To unsubscribe, e-mail: [email protected]
> > > > > For additional commands, e-mail: [email protected]
> > > > >
> > > > >
> > > > >
> > > >
> > > > Piergiorgio Lucidi
> > > > Mobile: 3395381669
> > > >
> > >
> >
> >
> > --
> > Piergiorgio
> >
>


-- 
Piergiorgio
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.