top of page

Graph based AI in the time of Astra and Fable - Combinatorial analysis of RAG, Search & Understanding information

Writer: Rajasankar Viswanathan
Rajasankar Viswanathan
5 hours ago
10 min read


Tl;dr

  1. Finding unknowns or Exploratory using LLMs is still problematic and verifying the output it not possible as cant set the 

  2. Enterprises still use RAG as a way to get the LLMs to generate correct information. 

  3. Using Embeddings for RAG misses context and results can't be verified. 

  4. Even if assume LLMs work perfectly still there are problems

  5. NaturalText AI is finding combinational solution bounds without prior data.

  6. NaturalText AI solves problems in current statistical clustering methods

  7. NaturalText AI finds connections in the data to make it network data and visualization. 

  8. NaturalText AI generates hierarchical clustering where solution bounds can be ranked with staggered confidence scores. 

  9. A fully deterministic method generates solutions when more than one answer for the problem exists. 

  10. Unlike LLMs the answer doesn't change how the question is asked. 

  11. Qualitative distribution is possible in text, tabular data, spreadsheets without any vectorization or embeddings. 

  12. Comparison with document handling methods, existing statistical clustering methods are discussed. 

  13. Answered the questions of why cant this be solved by AI agents or direct dumping of data into LLMs.


In the time of Astra and Fable why would anyone need to use another method that too an graph based AI that claims to find solutions better than LLMs? 


But people still use RAG, still use other databases etc because LLMs still need harnesses and supplementary information to do the tasks. 


So let us look at what are the needs of Enterprises, what exactly the NaturalText AI is and how it can do better than LLMs without being trained on large amounts of data. 


Before starting further, can’t this be done by someone to spin 100s of agents to do that work? 

Why does anyone need anything else when an agent swarm can do the work? 


The answer is obvious, LLM agents are using LLMs and even with all the advances LLMs are still not good for Enterprise tasks. There is one word to explain the problem, Exploratory. 


Most if not all of the tasks Enterprises do, is of exploratory nature. That can't be simply automated or even any probabilistic automation with the LLMs too. Let us see those in detail but an analogy in the order. 

Coding is the best example of good implementation of LLMs. Just type what you want into the chat and LLMs generate the coding. It can generate more code in an hour than a developer can write in a lifetime. Now, companies understood that coding is essentially exploratory and it needs more harness, more checks and more review than just generating code. 


Coding is a problem output is tentatively known, this task needs to be done repetitively, the data to be shown in this way and so on. However, people are trying to understand the information in documents, data formats etc wont even know what they are looking for. That needs more reading and more understanding. It can't simply be automated. 


Here it is not an argument on cost / speed / hallucinations. It assumes that LLMs work perfectly for the given task and without rate limits with unlimited tokens. Applying that will reduce the utility but discussing that in later. 


Now, let assume that LLMs work perfectly and do the comparison. 


Context - Reason for retrieval outside of LLM


Why still companies are using other retrieval methods outside of LLM has an answer. 

Context. 

Because LLMs may lack the necessary information to generate an answer, the companies need to provide correct background information to the LLMs along with the question to get a contextual answer. 


Embeddings and RAG 


How the extra context fed to LLM is the work of Retrieval Augmented Generation and statistical conversion of text called embeddings. 


Embeddings are a series of numbers that (supposedly) captures the essence of the document. Depending upon LLM and metric called dimension this series of numbers can be compared with each other to find whether the documents are similar or not. It just tells you about similarity, not anything else. Dimensions indicate how many numbers there, 1024 is a good dimension whereas 382 is also available. 


The workflow goes like this. 

  1. Get the embeddings of all the documents. 

  2. Store the embedding in a special database called vector database or frameworks

  3. Get the embeddings for the query. 

  4. Using the frameworks to find similar documents 

  5. Feed the documents to LLMs along with query to generate answers 


In other words, private document search has a fancy name now. That is RAG. Pre-LLM methods are simply word counts and dont work, however, this is not simple as it looks because finding a correct document and feeding it to LLM is still hard. 


As we have seen feeding the LLMs directly may not work, let me explain what is the alternative. 


Let us start with the fourth step of the process, ie finding similar documents or otherwise known as clustering. The correct name should be statistical clustering or numerical clustering because it uses cosine similarity calculations to find similarity. 


Why it is known as just clustering instead of statistical clustering is because it was the default method. As no other clustering methods are available people just used it. 


Problem with RAG

RAG uses similarity calculation frameworks such as FAISS from Facebook or extensions to existing databases such as pgvector to PostgreSQL/PostgreS. The issue is these methods use statistical clustering and generate top-k documents on every single query for feeding LLMs. Here top-k means the required number of documents denoted by k to be generated from these results. 

It generates similar documents but won't provide any overall understanding because the information is still in bits and pieces. 


Problem with statistical clustering methods 

There are only a few statistical clustering methods such as k-means clustering, these methods have various limitations such as only work on small size of data and takes long time to generate results. Though those limitations are to an extent can be overcome by using embeddings  or newer methods such as Super-k-means. 

Even if the speed and processing limitations are overcome, the basic issue is how it is clustered. Context and nuances has no place in statistical methods. 


Network graph and algorithms

Another question would be why cant simply dump the data into a graph database and let algorithms such as Louvain algorithm or Leiden algorithm to use the network effects for clustering or extracting pattern matching. 

The reason is simple. There won't be any connections. Data cant be simply dumped into a database even if it is a Relational database. The relationships need to be created. The only difference between relational database and graph database is design and schema. Relational databases need a prior designed schema whereas graph databases would be flexible. 

Without the relationships databases are as good as text files. 


Combinatorial methods by NaturalText AI. 

NaturalText AI method can be called combinatorial clustering, the correct nomenclature should be finding combinatorial bounds from data. 

Simply put it finds the combinations of parameters or values that make sense of data. 


In a thousands of documents what makes one document similar to another? This is the question answered by NaturalText AI. This can be used in any format of data such as bio-sequences (DNA,RNA), tabular data or data sits in spreadsheets, chemical fingerprints and so on. 


Solving exploratory and finding unknowns 

Data analysis and Document analysis is an exploratory task and finding unknowns from data. This is the reason why most of Enterprise or corporate work is not automated to begin with. 

In exploratory analysis, the solutions can’t be tested or understood because the answer is unknown to begin with. 


Hierarchical clustering for various levels of solution bounds

In statistical clustering, hierarchical clustering means just two step clustering. However, most of the problems need staggered solutions with various degrees of ranking score or confidence. The analyst wanted to know what would be the result if a 50% confidence score is accepted. In real terms it would be “what is the risk involved?” the answer could range from 0-100. 

NaturalText Graph based combinatorial clustering offers that kind of hierarchical clustering with staggered confidence scores. 


Power of Training and Data vs Power of Combinatorial maths

The power of LLMs and Embeddings are from entire internet data not from methods whereas the power of NaturalText AI from mathematics alone without needing data or training the data. 


Combinatorial distribution or Qualitative distribution 


There are two ways to look at the distribution of data/documents in this. Here the definition of document includes any free standing long form text - pdfs, html files, chat, social media post, code etc. 

Converting the data into graph for visualization

NaturalText AI finds connections between datapoints in any data and shows how it is connected. This makes it easier for the data to be stored in a graph database or see it in graph visualizations. 

More than one answer exists

The same question could have more than one answer in various domains. For example, why the people like this movie would get various answers because different people could like different parts of the movie. A single solution answer wouldn't work out even with all the reasoning. 



Sl no

Usecase

LLMs based Embeddings

NaturalText Graph AI

1

RAG

Cosine Similarity based methods, clustering top-k entries every query

Pre-populated clusters and staggered confidence scores. 

2

Clustering methods 

Statistical and non-hierarchical clustering

Graph based combinatorial clustering

3

Converting into Network graph

Needs an convoluted and all-against-all method to find connections

Out-of-box working on the Graph. Works on any data, any size, any format.

4

Solving exploratory and finding unknowns 

Very limited or none. 

Finds bounds using combinatorial analysis. 

5

Hierarchical clustering for various levels of solution bounds

Separate methods exist but only two stage hierarchical clustering. 

Hierarchical clustering is the default with layers decided upon automatically.

6

Power of Methods

Comes from data and compute

From mathematical models

7

Qualitative distribution 

Very limited. 

As good as statistical distribution. 

8

More than one answer exists

Finds similarity

Finds the combinatorial bounds


Distribution of Concepts


When we are dealing with text data, we are looking at the concepts, facts and other opinions. In that qualitative distribution as in statistical or quantitative distribution is important. The success of the task depends on the understanding of this spread in the data, especially text data. 

This is a business translation of dense and sparse matrices. 



SL No

Usecase

LLM based Embeddings

NaturalText Graph AI

1

Concepts are too widespread 

Both the methods works fine as there is no overlapping or narrow intersections


Both the methods works fine as there is no overlapping or narrow intersections


2

Concepts are too compact - dealing same narrow concepts


Needs to be split into chunks and generate embeddings because narrow concepts may not generate widely differentiating comparisons

NaturalText Graph AI works fine without any modification as it internally makes split and clusters it. 


3

Mixed some of data is narrow but others are widespread


Embeddings suffers from the same but in this case finding those differentiations is hard


NaturalText Graph AI works fine as it can clearly separate two groups more easily. 

4

Mixed languages - same topic but in different languages 


Embeddings - if languages are popular languages embeddings could work better but not for not so popular languages.

NaturalText Graph AI works good as it can clearly separate languages but cant find similarities between languages. 

5

Text with other data such as numbers, tabular data

Embeddings - May work at all as even the multimodal embeddings should have trained on the data. 

Graph AI works natively as data can be mapped out in the graph plane and similarities found


Distribution of Documents 

For the tasks, understanding what kind of spread is in the documents and how it can work with Emebedding is another element missing from the Enterprise discussions. Let see the use cases and how the Embeddings and NaturalText match. 



Sl no

Case

Embedding

NaturalText Graph AI

1

Find near duplicate or similar documents? 

Works great. It can find the near duplicates and similar documents easily via all - against all comparison with other frameworks. Hard for documents count exceeds 100k. 

Works great. It can find the near duplicates and similar documents easily via clustering directly for any data size. 

2

Tracking changes or document versioning

Same as the previous case, embeddings can't easily tell where exactly the change occurred but that can be solved by adding a few layers such as chunking comparison, Entity extraction. Adds more compute and time. 

Can show the differences directly as it compares each data point such as word rather than processing entire document

3

How two different documents compare on concepts? 

Can’t show that directly but adding a few layers of compute such as chunking

Shows that directly as it works on understanding concepts. 

4

Finding spam, filtering unwanted

This is a complex case for the embeddings as clustering and ranking would be hard. Clustering frameworks may solve this problem for a small number of documents with more complexity. 

Can show the spread and ranking i.e. what kind of documents and clusters are more popular. 

NSFW or spam documents would be neatly separated. 

5

Separating Multilingual documents

Same as previous cases however, may not work for complex cases where languages are not popular or mixed, written in different scripts i.e. one language written in another language script. Same problem as language detection. 

Can separate the documents because combinatorial clustering works on a fine grained level. 

6

Group documents based on topics

Grouping or Clustering has to take the advantage of frameworks such as SuperK-means or Chunkdot but fine grained or large numbers of documents would take longer and may not be possible. 

Default behaviour. Comes with hierarchical clustering. 

7

Documents into Graph

Use of specific frameworks and finding document similarity can be converted into. Using direct extraction of Entities via LLM it is possible to convert the documents into Graph. 

Direct conversion as it shows the cluster based interaction. Using alignment based learning, entities can be extracted easily. 



Discussing the existing vector frameworks and clustering methods

There are quite few methods for calculating and clustering using embeddings. The main frameworks are 

FAISS (Facebook AI Similarity Search) is an open-source library developed by Meta for efficient similarity search and clustering of dense vectors

Chunkdot is an open-source Python library for  Multi-threaded matrix multiplication and cosine similarity calculations for dense and sparse matrices. 

Super K-Means Super fast clustering for high-dimensional vectors on CPUs (x86, ARM) and GPUs — for Python and C++. Faster clustering of vector embeddings than FAISS


Sl no

Usecase

Clustering Methods

FAISS,Chunkdot,Super K-Means

NaturalText Graph AI

1

Clustering

Similarity, k-means

Graph based combinatorial clustering

2

All-against-all comparison of data

No 

Yes

3

Number of clusters

Pre-determined, this means overlaps and running clustering several times for correct or acceptable clusters

From the data the clusters are automatically extracted.

4

Data size/Scaling

Limited to small size, indexing and querying is different from clustering. 

Any size of data can be clustered. 

5

Mixed data

Depends upon embeddings, some data cant be used because embeddings wont be generated or useless.

Works for pure text, long form and mixed data

6

Compute

High memory and processor power needed. 

Works on Low memory and compute. 

7

Hierarchical Clustering

No

Yes

8

Speed

Depends up on the datasize

Fastest on any datasize


 
 
 

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
NaturalText_Logo-09.png
About

NaturalText AI uses novel artificial intelligence technology to uncover patterns and reveal insights hidden in text data.

NaturalText, Inc.

Delaware, USA

Navigation
NaturalText_Website-Buttons_Request-a-De
  • Instagram
  • Facebook
  • Twitter
  • LinkedIn
  • YouTube
  • TikTok
Contact

Thank you for submitting! We will be in touch.

© 2025 NaturalText, Inc. All rights reserved.

bottom of page