https://linen.dev logo
Join Slack
Powered by
# good-reads
  • h

    Hugo Lu

    07/10/2025, 8:04 AM
    https://tinyurl.com/4mwve2yk
  • c

    Christopher Bergh

    07/24/2025, 1:24 PM
    https://datakitchen.io/fitt-data-architecture/
    👀 1
  • k

    Kumari Surya Remanan

    08/04/2025, 9:21 AM
    Hi everyone — I just submitted a GitHub Issue proposing a new blog topic: "Leveraging EV Infrastructure for Real-Time Data Synchronization Between Charging Stations and AI Models Using Airbyte" Here’s the link: https://github.com/airbytehq/airbyte/issues/64485 Would love to get feedback and see if it aligns with the community content plan!
    👍 1
  • b

    Bala

    09/09/2025, 11:28 AM
    Hello team, I'd submitted an article for review through the form about a month ago. Could you please let me know how long the review process generally takes?
    👍 1
  • u

    سیدحماد احمد

    09/18/2025, 6:27 AM
    I have also submitted an article (What's New in dbt 1.10) for review through the form but haven't heard back. Could you please confirm if you are reviewing the submission?
  • c

    Chiara

    10/23/2025, 1:12 PM
    Hi everyone! I am organizing OSA Con 2025 and it's just around the corner. Lots of great speakers this year: Amazon, Snowflake, Preset, Percona, TiDB, Apple, Nutanix, Altinity, and more! 🗓️ Nov 4-5 📍ONLINE Register here: https://osacon.io/
  • v

    Vivek Dubey

    10/30/2025, 4:31 AM
    💠 dbt Coalesce 2025: What 14,000 Practitioners Learned 📃 Beyond tools and trends, Coalesce 2025 revealed the blueprint for data systems that can think, learn, and earn trust. ✉️ Read the complete article here: https://metadataweekly.substack.com/p/dbt-coalesce-2025-what-14000-practitioners-learned
  • y

    Young

    10/30/2025, 2:29 PM
    🎊🎊🎊 Join us on Nov 10 in San Francisco for the next Data for AI Meetup! 🎟️ RSVP: https://luma.com/p7m6mxki?locale=en The Agentic AI era is here, and data stacks need to catch up. Analytics used to be all customer-facing, now it's agent-facing, too. Join the event and see how multi-modal catalogs and real-time lakehouses can help you build a data infra that's ready for AI and agents. Hear from top engineers at Uber, Pinterest, Datastrato, and VeloDB(https://www.velodb.io/): - Lessons from Uber: Re-architecting metadata systems to manage 200B+ entries. - From hours to seconds: How Pinterest solved its 130 petabytes data partition listing challenge. - Design AI-native, agent-driven data operations with Gravitino. - Build a real-time lakehouse on AWS with Glue, S3 Tables, and Apache Doris to get an AI-ready data foundation. 🍺 Drinks, food, and great conversations guaranteed 💬 Nov. 10, 17:30 - 20:30 PT @ AWS Builder Loft, San Francisco 🎟️ RSVP link at the top Grateful to our partners at AWS and Datastrato for supporting the event 🙏
  • v

    Vivek Dubey

    11/04/2025, 11:46 AM
    👀 The Semantic Gap: Why Your AI Still Can’t Read The Room Brilliant piece by Vince Dacanay, Head of Data at Prodege, LLC, on Metadata Weekly: > “Your AI can process a decade of data in seconds, but it still misses the point. It doesn’t catch the hesitation before an answer, or the unspoken politics behind a ‘yes.’” Vince calls this semantic density, the human layer of meaning that no AI can yet read. It’s the culture, context, and intuition that sit between words and understanding. He makes a sharp point: the best AI systems aren’t built to “understand everything.” They’re built with constraints — because the semantic gap isn’t a flaw, it’s where humans still matter most. It shows up when: → Questions carry unspoken stakes. → Experience compresses into two words. → Context fills in what data can’t see. How to close this gap??? Read the article for complete details here: https://metadataweekly.substack.com/p/the-semantic-gap-why-your-ai-still-cant-read-the-room
  • y

    Young

    11/04/2025, 2:32 PM
    🎟️*Apache Doris(https://doris.apache.org/)* Summit 2025 kicks off (only 7 hours left!) Join us for a full day of sessions featuring the latest innovations, user stories, and ecosystem insights around real-time analytics and search in the AI era. Check out the full agenda and register here👇 https://lnkd.in/gECm7n4E
  • v

    Vivek Dubey

    11/06/2025, 1:10 PM
    👑 The AI Era Runs on Context: "In the internet era, content was king. In the AI era, context is sovereign." At Re:govern 2025 - 20+ of the world’s most AI-forward data teams — from Workday, Mastercard, CME Group, Dropbox, GitLab, and more — dropped real talk on what’s working (and what’s not) in the age of AI + governance.
    One takeaway stood above all: Context is king.
    AI readiness starts with context readiness, and the best teams are building it before the crisis hits. Missed the live action? Catch every session + recap here → https://atlan.com/regovern/?utm_medium=outreach&utm_source=slack&utm_campaign=regovern_2025
  • y

    Young

    11/06/2025, 1:44 PM
    Webinar: Query Billions of JSON Rows and 10K+ Subcolumns in Seconds 👉 Register: https://lnkd.in/dPmtxRMf JSON is everywhere: logs, metrics, e-commerce, IoT, but most systems still struggle to query it efficiently at scale. Join our webinar to see how VARIANT in Apache Doris delivers fast, schema-flexible JSON analytics at scale: 1️⃣ Understand the landscape: How Elasticsearch, Snowflake, ClickHouse, and Iceberg handle JSON. And where Apache Doris stands out. 2️⃣ See how Apache Doris does it: Sparse columns, subcolumn compaction, and schema templates for performance and flexibility. 3️⃣ Watch the demo: Query 1 billion rows of JSON data in seconds on AWS deployment. No tricks. Just high-performance JSON analytics that scale. 📅 Nov. 20, 4:00 p.m. PT | 7:00 p.m. ET
  • v

    Vivek Dubey

    11/19/2025, 6:59 AM
    💭 What if your AI Analyst understood your business as well as your best human analyst? Yes, its the DREAM... but it needs a ton of Context Engineering. That’s the question Shubham Bhargav explores in the latest Metadata Weekly edition and IMO if there's only one article you can read this week, read this one! We’ve seen time and again that it’s not the models holding AI back, but the context gap — all the meaning that lives in people’s heads, not in systems. The definitions, judgment calls, and patterns of reasoning that make decisions make sense. We have been rolling up our sleeves getting these AI agents into production. Shubham breaks down what it actually takes to close the context engineering gap. He dives into how data teams can build a “context supply chain,” layer semantics, and continuously refine meaning through human–AI feedback loops. Read the complete article here: https://metadataweekly.substack.com/p/context-engineering-for-ai-analysts
  • y

    Young

    11/20/2025, 11:01 PM
    Hey all, we're Apache Doris (https://doris.apache.org/)and our JSON analytics and VARIANT webinar starts in an hour. 👉 Join us live in an hour: https://us06web.zoom.us/j/89475839940?pwd=2edKgJFO8QDEOnE55hMc4ByDIDIC1F.1 If you work with logs, events, IoT data, or any large-scale JSON workloads, this session will give you a practical breakdown of how different systems handle JSON. We’ll walk through: 1. How semi-structured analytics evolved: From TEXT and JSON to VARIANT 2. How major systems approach JSON: Apache Doris, Elasticsearch, Snowflake, ClickHouse, and Iceberg 3. How Apache Doris VARIANT type works: sparse columns, subcolumn vertical compaction, and schema templates 4. Live demo: Querying 1B rows of JSON in seconds on AWS using Apache Doris
  • v

    Vivek Dubey

    11/26/2025, 7:20 AM
    💎 Data can look “healthy”… and your model can still drift, hallucinate, or amplify bias. That’s the observability gap that Mahdi Karabiben unpacks in his latest article Metadata Weekly. We’ve spent years solving data observability for dashboards. But in the AI era, a clean pipeline doesn’t guarantee a safe decision — because the real risk now lives at the intersection of data + model + agent behavior. Mahdi breaks down why unified Data + AI Observability is quickly becoming essential for trustworthy AI systems, covering: ➡️ Why good data can still create bad AI ➡️ Why alerts need to be built for agents, not humans ➡️ How lineage becomes the control plane ➡️ Why the future is about decision trust, not just data trust If you’re aiming to build trustworthy AI systems, this one is worth the read. Read the full article on Metadata Weekly: https://metadataweekly.substack.com/p/data-trust-to-decision-trust-the
  • v

    Vivek Dubey

    12/01/2025, 6:57 AM
    Customer data is the highest-risk, highest-impact fuel for AI, and the easiest to get wrong. One incorrect field or outdated attribute can ripple through personalization, scoring, and support workflows within seconds. In the latest Metadata Weekly edition, Michele Nieberding digs into what it really takes to build AI agents you can trust with customer data. She breaks down the two foundations teams can’t ignore: ➡️ Governance — ensuring data is accurate, compliant, and purpose-aligned ➡️ Context — giving AI a deep semantic understanding of what the data means Her article covers everything from lineage as a non-negotiable to purpose-aware access, data minimization, and how to build the intelligence layer AI actually needs. If you’re building customer-data AI or experimenting with agents, you’ll want to read this one. ✉️ Read the full article on Metadata Weekly: https://metadataweekly.substack.com/p/building-ai-agents-you-can-trust
  • y

    Young

    12/09/2025, 10:54 AM
    Modern systems generate massive amount of JSON in logs, traces, device tags, custom properties...Apache Doris’s VARIANT type was built to handle JSON at scale, especially when JSON evolves fast, contains thousands of fields, and needs to be queried analytically. Three real-world use cases of JSON and VARIANT from our recent webinar: 1️⃣ Observability: Logs & Traces - Handles field type changes without data loss - Automatically adds/removes fields from the schema - Achieves 5X compression vs storing JSON as strings - Huge savings on storage for log-heavy workloads 2️⃣ IoT & IoV - Apache Doris delivers better analytical performance than traditional time-series databases - Can run nested/array functions directly (e.g., array_contains(tags['b'], 1)) 3️⃣ SaaS Companies - Supports thousands of dynamic customer-defined fields - Rich indexing for high concurrency customer-facing dashboards - Full-text search + JOINs in the same engine. Hard to find outside Apache Doris. Grateful to the insightful sharing from Owen Xiao, Apache Doris PMC Member and VeloDB Product VP. 👉 Watch the webinar in full: https://lnkd.in/g6JKcx4w
  • y

    Young

    12/15/2025, 2:56 PM
    ⏰ Last Call: Don’t miss our webinar on Hybrid Search & Context Engineering with Apache Doris 4.0 🎟️ RSVP: https://lnkd.in/gvczv4zi AI models won't work well without the right context. But retrieving the right context at scale (think datasets in the billions) is a real challenge. Join us to learn how Apache Doris 4.0 solves this with Hybrid Search, bringing together: 1️⃣ Vector search for semantic relevance 2️⃣ Full-text search (BM25 + inverted index) 3️⃣ High-performance structured analytics All in one unified SQL engine Come and see real use cases of how data teams are using Doris 4.0 to build accurate and cost efficient their AI application.
  • v

    Vivek Dubey

    12/16/2025, 12:19 PM
    ✈️ 𝗪𝗲’𝘃𝗲 𝗵𝗶𝘁 𝘁𝗵𝗲 𝗰𝗲𝗶𝗹𝗶𝗻𝗴 𝗼𝗳 𝘄𝗵𝗮𝘁 𝗔𝗜 𝗰𝗮𝗻 𝗱𝗼 𝘄𝗶𝘁𝗵𝗼𝘂𝘁 𝗱𝗲𝗲𝗽 𝗲𝗻𝘁𝗲𝗿𝗽𝗿𝗶𝘀𝗲 𝗰𝗼𝗻𝘁𝗲𝘅𝘁. That was the signal I and many couldn’t ignore at AWS re:Invent 2025. This year wasn’t really about bigger models or faster chips. It was about finally acknowledging what practitioners have been saying for months: "intelligence doesn’t come from models alone. It comes from the context that systems can understand." In this latest edition of community led Metadata Weekly, guest author shared a few patterns that stood out at re:Invent 2025: ➡️ The shift from retrieving context → to embedding it → to embodying it across training, memory, and infrastructure ➡️ Why voice agents, code assistants, and text-to-SQL all break at the same place: missing metadata ➡️ How Nova Forge, Frontier Agents, AI Factories, and Transform Custom are actually context systems in disguise ➡️ Why definitions, lineage, and business rules are becoming the true enterprise moat, not model size Models will commoditize. Compute will get cheaper. But the companies that can create, govern, and distribute high-quality context will build the systems the world trusts. 📥 𝗥𝗲𝗮𝗱 𝘁𝗵𝗲 𝗳𝘂𝗹𝗹 𝗲𝗱𝗶𝘁𝗶𝗼𝗻 𝗼𝗻 𝗠𝗲𝘁𝗮𝗱𝗮𝘁𝗮 𝗪𝗲𝗲𝗸𝗹𝘆: https://metadataweekly.substack.com/p/aws-reinvent-2025-what-reinvent-quietly
  • v

    Vivek Dubey

    02/17/2026, 11:11 AM
    We knew foundations mattered. We just hoped they wouldn’t matter this time. That line from Gaurav Ramesh at OpenTable captures what 2025 actually revealed about AI. Most AI initiatives didn't fail because the models weren't ready. They failed because our foundations weren't built for what we asked them to do. Frontier models were positioned as general-purpose intelligence. Powerful enough to work around messy data, unclear ownership, and brittle systems. Most leaders knew foundations mattered. We just hoped the models would let us postpone the hard work. That bet didn’t pay off. A refreshing way to think about 2025's failures? Stalled pilots aren't dead ends. They are the cost of learning. They expose real constraints and clarify what actually needs to change. And budget isn't the missing piece. Money doesn't create clarity — organizational confidence does. The missing foundation isn't data quality or governance. It's self-awareness. AI doesn’t just test infrastructure. It tests how well we understand how we work. 👇 Read Gaurav’s full piece in Metadata Weekly: https://metadataweekly.substack.com/p/the-human-elements-of-the-ai-foundations
  • j

    J. Riso

    02/20/2026, 6:39 PM
    I did an analysis of what LLMs think Airbyte (Cloud) costs vs. two other ETL products for a generic use case. Dunno if it's a "good read" but some folks here might be interested. Also shared on LinkedIn https://risogroup.co/insights/llm-etl-pricing.html
  • v

    Vivek Dubey

    02/26/2026, 11:18 AM
    𝟯𝟴% 𝗯𝗲𝘁𝘁𝗲𝗿 𝗔𝗜 𝗮𝗰𝗰𝘂𝗿𝗮𝗰𝘆. 𝗡𝗼 𝗻𝗲𝘄 𝗺𝗼𝗱𝗲𝗹. 𝗡𝗼 𝗻𝗲𝘄 𝗱𝗮𝘁𝗮. 𝗝𝘂𝘀𝘁 𝗯𝗲𝘁𝘁𝗲𝗿 𝗰𝗼𝗻𝘁𝗲𝘅𝘁. That’s the headline from a controlled NL-to-SQL experiment discussed by Manoj Shanmugasundaram in Metadata Weekly. Across 522 query evaluations, the only variable that changed was context quality — and it made all the difference. Concise, high-signal context (business definitions, SQL patterns, domain rules) drove a 38% accuracy gain. Verbose, catalog-style documentation? Performance dropped and costs rose. More words diluted the signal. The biggest lift wasn't on simple or extreme queries. It was on medium-complexity ones — the joins and aggregations that make up everyday analytics work — where focused context delivered a 𝟮.𝟭𝟱𝘅 𝗶𝗺𝗽𝗿𝗼𝘃𝗲𝗺𝗲𝗻𝘁. The mindset shift: metadata was built for humans to browse. Now it also needs to work for machines to reason. Most teams are still optimizing for readability, not machine usability. If your "talk to data" initiative is stalling, it might not be a model problem. It might be a context problem. Manoj breaks down what machine-usable context actually looks like — and how to get started without rebuilding your stack. Read it in Metadata Weekly: https://metadataweekly.substack.com/p/the-context-problem-nobodys-fixing
  • v

    Vivek Dubey

    03/18/2026, 3:51 PM
    A pricing agent goes rogue. Everyone blames the model. Nobody governed the context it was reasoning over. Most AI governance failures aren't exotic model problems. They're old governance failures that AI makes visible. At Gartner D&A Summit this week, I kept hearing the same debate: data governance vs. AI governance. Two tracks. Two teams. IMO, it's the wrong frame. Gartner's agenda said as much: a unified "Data and AI Governance" track. Not two separate sessions. One. 42% of enterprises already have AI agents in production. The governance frameworks haven't caught up. The disciplines aren't new. Data quality, lineage, access control, lifecycle management. What changed is the surface area. Those practices now need to cover models, prompts, retrieval pipelines, and autonomous agents. One thing. Not three. Charlotte Ledoux and Vivek Dubey make the case for what that actually looks like. Worth the read: https://metadataweekly.substack.com/p/data-governance-vs-ai-governance
  • t

    Tim OBrien

    05/21/2026, 2:22 PM
    I just published on amazon about GenAI and multi-agent systems. I included an entire chapter on Airbyte. Here is the link: https://www.amazon.com/Building-Complex-Multi-Agent-Systems-Prompting-ebook/dp/B0GZW71GJH/ref=cm_cr_arp_d_product_top?ie=UTF8
  • v

    Vivek Dubey

    06/01/2026, 7:50 PM
    The most asked question in enterprise AI right now: "What actually is a context layer?" Everyone uses the term. Almost no one defines it the same way. Prukalpa breaks it down clearly: the 3 substrates that form machine-usable context and the 5 capabilities that build an enterprise context layer. A context layer turns three things into machine-usable context for AI: → Knowledge — what the business means → Expertise — how work actually gets done → Norms — what's allowed This is why agents dazzle in demos and break in production. Most architectures have knowledge. They're missing expertise and norms. Full breakdown in this week's Context & Chaos: https://metadataweekly.substack.com/p/what-an-enterprise-context-layer
  • n

    Nitin Jain

    06/24/2026, 10:44 PM
    Snowflake summit update : https://cdatainsights.com/blogs/snowflake-summit-2026-recap
  • v

    Vivek Dubey

    06/29/2026, 10:58 AM
    ❄️ Snowflake Summit 2026: Everyone Owns Context Now Snowflake Summit could have been a context drinking game. Every vendor claimed it. They all meant something different. Four days at Summit. Semantic models, metadata connectors, agent governance, memory. But each one defined context as the slice its own product happens to produce. Semantics for the BI tools, schemas for the warehouses, embeddings for the vector stores. If context only means what a platform already does, it was never about agents understanding the business. It was about the vendor's center of gravity. The stage showed the technical half. What attendees kept asking about was the organizational half. Who maintains definitions when the team that built them turns over? When two business units define the same metric differently and an agent has to answer a question that spans both, who decides? Where does that decision live so every agent can act on it? You'll have hundreds of agents and they won't live in one place. With how fast models are changing and new platforms popping up, you can't afford to lock your context inside whichever tool a team started in. This article went deep on what was announced, what customers were actually asking, and the work that's still ahead: https://contextandchaos.substack.com/p/snowflake-summit-2026-everyone-owns
  • t

    thiago

    06/30/2026, 2:22 PM
    Introducing AI-Lake Format 🏔️ We're open-sourcing ailake — a self-contained file format for AI Lakehouses, written in Rust. The problem: Running BI and GenAI on the same data means juggling two stacks — a Data Lake for analytics and a separate vector DB (Pinecone, Milvus) for RAG. Two systems, two sources of truth, double the ops cost. Our answer: One file. Tabular data + embeddings + HNSW index — all inside a single extended Parquet file on S3. Key properties: - 100% Iceberg-compatible — Spark, Trino, DuckDB, PyIceberg read it natively with zero plugins - Vector search built-in — HNSW index lives in the file footer; loaded via mmap, never pulled into RAM whole - Geometric pruning — query vector compared to per-file centroids at manifest scan time; eliminates 95–99% of files before any S3 I/O - Hybrid search — BM25 + HNSW + Tantivy FTS in the same query pipeline - ACID via Iceberg — snapshots, time-travel, compaction, deletes — all standard - Python-first — import ailake, zero native deps from user side (PyO3 wheel) No new infra. Your existing Iceberg catalog, your existing S3 bucket. Add vector search without moving your data. GitHub: github.com/ailake-io/ai-lakehouse Would love feedback — especially from anyone running RAG at scale on a Lakehouse.
  • t

    thiago

    07/14/2026, 6:03 PM
    🚀 The Future of Data for AI: AI-Lakehouse + DuckLake in v0.1.4! Data integration for Artificial Intelligence just took a massive leap forward. The v0.1.4 release of ai-lakehouse now features native support for DuckLake! Since the core mission of ai-lakehouse is to bridge analytical data storage with the power of vectors (embeddings) and hybrid search for AI, the addition of DuckLake as a catalog backend is a total game-changer. 🧠 Vector Power Meets DuckLake Simplicity: Native & Ultra-Fast Vector Search: ai-lakehouse extends the analytical ecosystem to store vector representations of text, images, and audio. With DuckLake, you manage these embeddings and metadata directly within a lightweight SQL database (like DuckDB), while keeping the raw data files in Parquet. Hassle-Free Hybrid Search (FTS + Vectors): It is now incredibly simple to run hybrid queries combining traditional Full-Text Search (keyword-based) with Dense Vector Search (semantic similarity). All of this without having to manage a separate, complex vector database. Efficient, Local RAG Pipelines: If you are building Retrieval-Augmented Generation (RAG) applications for LLMs, this combination removes all infrastructure friction. You get to store your documents, metadata, and embeddings in an ACID-compliant, portable lakehouse with top-tier performance. Vector Versioning with Time Travel: Because DuckLake supports version control and Time Travel, you can easily track how your vector space evolved over time. This is critical for auditing AI models and ensuring experiment reproducibility. Zero-Cost, Serverless Vector Architecture: Forget paying for and maintaining expensive vector database clusters. The combined architecture of AI-Lakehouse and DuckLake runs serverless and locally, scaling seamlessly from your laptop to cloud object storage (S3/ADLS/GCS). The ultimate synergy of a flexible data lake, a transactional data warehouse, and a local vector database. 👉 Check out the v0.1.4 release notes and start building the next generation of AI applications today! github.com/ailake-io/ai-lakehouse
  • v

    Vivek Dubey

    08/10/2026, 5:32 PM
    Gartner expects the average global Fortune 500 to run more than 150,000 AI agents by 2028, up from fewer than 15 in 2025. When your exec asks for the list, don't build it. Read this week's Context & Chaos on why, and on what to do instead. A list kept by hand is wrong before it is finished, and most of the people building agents never filed a ticket to do it. So govern one level down. An agent is a composition: skills, tools, MCP servers, credentials, memory. Those parts sit still. The agent does not. It gets renamed, forked, cloned by someone who did not know the first one existed. Ownership assigned to an agent expires the moment someone renames it. Ownership assigned to a skill survives every agent built on it. Someone will say this is just source control with extra steps. It is not. Git holds versions of files. It cannot tell you which five agents depend on this skill, or what it costs to run. Those registry questions are where a lot of governance programs will fail next year. Pick one skill more than one of your agents uses. Ask who owns it, which agents depend on it, and what breaks if you delete it tonight. More than five minutes, and the skill is not governed. Neither is anything built on it. Read the full piece on Context & Chaos: contextandchaos.substack.com/p/your-agents-are-code-stop-governing