Skip to content

Add HTML/Webpage Parsing Layer for RAG Pipeline #82

Description

@aliamerj

Right now, our RAG pipeline only handles PDFs, but a ton of valuable content lives on web pages and in raw HTML. extend our parser to ingest HTML documents directly and pull out both text and visuals.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions