Graph based AI in the time of Astra and Fable - Combinatorial analysis of RAG, Search & Understanding information

Tl;dr
Finding unknowns or Exploratory using LLMs is still problematic and verifying the output it not possible as cant set the
Enterprises still use RAG as a way to get the LLMs to generate correct information.
Using Embeddings for RAG misses context and results can't be verified.
Even if assume LLMs work perfectly still there are problems
NaturalText AI is finding combinational solution bounds without prior data.
NaturalText AI solves problems in current statistical clustering methods
NaturalText AI finds connections in the data to make it network data and visualization.
NaturalText AI generates hierarchical clustering where solution bounds can be ranked with staggered confidence scores.
A fully deterministic method generates solutions when more than one answer for the problem exists.
Unlike LLMs the answer doesn't change how the question is asked.
Qualitative distribution is possible in text, tabular data, spreadsheets without any vectorization or embeddings.
Comparison with document handling methods, existing statistical clustering methods are discussed.
Answered the questions of why cant this be solved by AI agents or direct dumping of data into LLMs.
In the time of Astra and Fable why would anyone need to use another method that too an graph based AI that claims to find solutions better than LLMs?
But people still use RAG, still use other databases etc because LLMs still need harnesses and supplementary information to do the tasks.
So let us look at what are the needs of Enterprises, what exactly the NaturalText AI is and how it can do better than LLMs without being trained on large amounts of data.
Before starting further, can’t this be done by someone to spin 100s of agents to do that work?
Why does anyone need anything else when an agent swarm can do the work?
The answer is obvious, LLM agents are using LLMs and even with all the advances LLMs are still not good for Enterprise tasks. There is one word to explain the problem, Exploratory.
Most if not all of the tasks Enterprises do, is of exploratory nature. That can't be simply automated or even any probabilistic automation with the LLMs too. Let us see those in detail but an analogy in the order.
Coding is the best example of good implementation of LLMs. Just type what you want into the chat and LLMs generate the coding. It can generate more code in an hour than a developer can write in a lifetime. Now, companies understood that coding is essentially exploratory and it needs more harness, more checks and more review than just generating code.
Coding is a problem output is tentatively known, this task needs to be done repetitively, the data to be shown in this way and so on. However, people are trying to understand the information in documents, data formats etc wont even know what they are looking for. That needs more reading and more understanding. It can't simply be automated.
Here it is not an argument on cost / speed / hallucinations. It assumes that LLMs work perfectly for the given task and without rate limits with unlimited tokens. Applying that will reduce the utility but discussing that in later.
Now, let assume that LLMs work perfectly and do the comparison.
Context - Reason for retrieval outside of LLM
Why still companies are using other retrieval methods outside of LLM has an answer.
Context.
Because LLMs may lack the necessary information to generate an answer, the companies need to provide correct background information to the LLMs along with the question to get a contextual answer.
Embeddings and RAG
How the extra context fed to LLM is the work of Retrieval Augmented Generation and statistical conversion of text called embeddings.
Embeddings are a series of numbers that (supposedly) captures the essence of the document. Depending upon LLM and metric called dimension this series of numbers can be compared with each other to find whether the documents are similar or not. It just tells you about similarity, not anything else. Dimensions indicate how many numbers there, 1024 is a good dimension whereas 382 is also available.
The workflow goes like this.
Get the embeddings of all the documents.
Store the embedding in a special database called vector database or frameworks
Get the embeddings for the query.
Using the frameworks to find similar documents
Feed the documents to LLMs along with query to generate answers
In other words, private document search has a fancy name now. That is RAG. Pre-LLM methods are simply word counts and dont work, however, this is not simple as it looks because finding a correct document and feeding it to LLM is still hard.
As we have seen feeding the LLMs directly may not work, let me explain what is the alternative.
Let us start with the fourth step of the process, ie finding similar documents or otherwise known as clustering. The correct name should be statistical clustering or numerical clustering because it uses cosine similarity calculations to find similarity.
Why it is known as just clustering instead of statistical clustering is because it was the default method. As no other clustering methods are available people just used it.
Problem with RAG
RAG uses similarity calculation frameworks such as FAISS from Facebook or extensions to existing databases such as pgvector to PostgreSQL/PostgreS. The issue is these methods use statistical clustering and generate top-k documents on every single query for feeding LLMs. Here top-k means the required number of documents denoted by k to be generated from these results.
It generates similar documents but won't provide any overall understanding because the information is still in bits and pieces.
Problem with statistical clustering methods
There are only a few statistical clustering methods such as k-means clustering, these methods have various limitations such as only work on small size of data and takes long time to generate results. Though those limitations are to an extent can be overcome by using embeddings or newer methods such as Super-k-means.
Even if the speed and processing limitations are overcome, the basic issue is how it is clustered. Context and nuances has no place in statistical methods.
Network graph and algorithms
Another question would be why cant simply dump the data into a graph database and let algorithms such as Louvain algorithm or Leiden algorithm to use the network effects for clustering or extracting pattern matching.
The reason is simple. There won't be any connections. Data cant be simply dumped into a database even if it is a Relational database. The relationships need to be created. The only difference between relational database and graph database is design and schema. Relational databases need a prior designed schema whereas graph databases would be flexible.
Without the relationships databases are as good as text files.
Combinatorial methods by NaturalText AI.
NaturalText AI method can be called combinatorial clustering, the correct nomenclature should be finding combinatorial bounds from data.
Simply put it finds the combinations of parameters or values that make sense of data.
In a thousands of documents what makes one document similar to another? This is the question answered by NaturalText AI. This can be used in any format of data such as bio-sequences (DNA,RNA), tabular data or data sits in spreadsheets, chemical fingerprints and so on.
Solving exploratory and finding unknowns
Data analysis and Document analysis is an exploratory task and finding unknowns from data. This is the reason why most of Enterprise or corporate work is not automated to begin with.
In exploratory analysis, the solutions can’t be tested or understood because the answer is unknown to begin with.
Hierarchical clustering for various levels of solution bounds
In statistical clustering, hierarchical clustering means just two step clustering. However, most of the problems need staggered solutions with various degrees of ranking score or confidence. The analyst wanted to know what would be the result if a 50% confidence score is accepted. In real terms it would be “what is the risk involved?” the answer could range from 0-100.
NaturalText Graph based combinatorial clustering offers that kind of hierarchical clustering with staggered confidence scores.
Power of Training and Data vs Power of Combinatorial maths
The power of LLMs and Embeddings are from entire internet data not from methods whereas the power of NaturalText AI from mathematics alone without needing data or training the data.
Combinatorial distribution or Qualitative distribution
There are two ways to look at the distribution of data/documents in this. Here the definition of document includes any free standing long form text - pdfs, html files, chat, social media post, code etc.
Converting the data into graph for visualization
NaturalText AI finds connections between datapoints in any data and shows how it is connected. This makes it easier for the data to be stored in a graph database or see it in graph visualizations.
More than one answer exists
The same question could have more than one answer in various domains. For example, why the people like this movie would get various answers because different people could like different parts of the movie. A single solution answer wouldn't work out even with all the reasoning.
Sl no | Usecase | LLMs based Embeddings | NaturalText Graph AI |
1 | RAG | Cosine Similarity based methods, clustering top-k entries every query | Pre-populated clusters and staggered confidence scores. |
2 | Clustering methods | Statistical and non-hierarchical clustering | Graph based combinatorial clustering |
3 | Converting into Network graph | Needs an convoluted and all-against-all method to find connections | Out-of-box working on the Graph. Works on any data, any size, any format. |
4 | Solving exploratory and finding unknowns | Very limited or none. | Finds bounds using combinatorial analysis. |
5 | Hierarchical clustering for various levels of solution bounds | Separate methods exist but only two stage hierarchical clustering. | Hierarchical clustering is the default with layers decided upon automatically. |
6 | Power of Methods | Comes from data and compute | From mathematical models |
7 | Qualitative distribution | Very limited. | As good as statistical distribution. |
8 | More than one answer exists | Finds similarity | Finds the combinatorial bounds |
Distribution of Concepts
When we are dealing with text data, we are looking at the concepts, facts and other opinions. In that qualitative distribution as in statistical or quantitative distribution is important. The success of the task depends on the understanding of this spread in the data, especially text data.
This is a business translation of dense and sparse matrices.
SL No | Usecase | LLM based Embeddings | NaturalText Graph AI |
1 | Concepts are too widespread | Both the methods works fine as there is no overlapping or narrow intersections | Both the methods works fine as there is no overlapping or narrow intersections |
2 | Concepts are too compact - dealing same narrow concepts | Needs to be split into chunks and generate embeddings because narrow concepts may not generate widely differentiating comparisons | NaturalText Graph AI works fine without any modification as it internally makes split and clusters it. |
3 | Mixed some of data is narrow but others are widespread | Embeddings suffers from the same but in this case finding those differentiations is hard | NaturalText Graph AI works fine as it can clearly separate two groups more easily. |
4 | Mixed languages - same topic but in different languages | Embeddings - if languages are popular languages embeddings could work better but not for not so popular languages. | NaturalText Graph AI works good as it can clearly separate languages but cant find similarities between languages. |
5 | Text with other data such as numbers, tabular data | Embeddings - May work at all as even the multimodal embeddings should have trained on the data. | Graph AI works natively as data can be mapped out in the graph plane and similarities found |
Distribution of Documents
For the tasks, understanding what kind of spread is in the documents and how it can work with Emebedding is another element missing from the Enterprise discussions. Let see the use cases and how the Embeddings and NaturalText match.
Sl no | Case | Embedding | NaturalText Graph AI |
1 | Find near duplicate or similar documents? | Works great. It can find the near duplicates and similar documents easily via all - against all comparison with other frameworks. Hard for documents count exceeds 100k. | Works great. It can find the near duplicates and similar documents easily via clustering directly for any data size. |
2 | Tracking changes or document versioning | Same as the previous case, embeddings can't easily tell where exactly the change occurred but that can be solved by adding a few layers such as chunking comparison, Entity extraction. Adds more compute and time. | Can show the differences directly as it compares each data point such as word rather than processing entire document |
3 | How two different documents compare on concepts? | Can’t show that directly but adding a few layers of compute such as chunking | Shows that directly as it works on understanding concepts. |
4 | Finding spam, filtering unwanted | This is a complex case for the embeddings as clustering and ranking would be hard. Clustering frameworks may solve this problem for a small number of documents with more complexity. | Can show the spread and ranking i.e. what kind of documents and clusters are more popular. NSFW or spam documents would be neatly separated. |
5 | Separating Multilingual documents | Same as previous cases however, may not work for complex cases where languages are not popular or mixed, written in different scripts i.e. one language written in another language script. Same problem as language detection. | Can separate the documents because combinatorial clustering works on a fine grained level. |
6 | Group documents based on topics | Grouping or Clustering has to take the advantage of frameworks such as SuperK-means or Chunkdot but fine grained or large numbers of documents would take longer and may not be possible. | Default behaviour. Comes with hierarchical clustering. |
7 | Documents into Graph | Use of specific frameworks and finding document similarity can be converted into. Using direct extraction of Entities via LLM it is possible to convert the documents into Graph. | Direct conversion as it shows the cluster based interaction. Using alignment based learning, entities can be extracted easily. |
Discussing the existing vector frameworks and clustering methods
There are quite few methods for calculating and clustering using embeddings. The main frameworks are
FAISS (Facebook AI Similarity Search) is an open-source library developed by Meta for efficient similarity search and clustering of dense vectors
Chunkdot is an open-source Python library for Multi-threaded matrix multiplication and cosine similarity calculations for dense and sparse matrices.
Super K-Means Super fast clustering for high-dimensional vectors on CPUs (x86, ARM) and GPUs — for Python and C++. Faster clustering of vector embeddings than FAISS
Sl no | Usecase | Clustering Methods FAISS,Chunkdot,Super K-Means | NaturalText Graph AI |
1 | Clustering | Similarity, k-means | Graph based combinatorial clustering |
2 | All-against-all comparison of data | No | Yes |
3 | Number of clusters | Pre-determined, this means overlaps and running clustering several times for correct or acceptable clusters | From the data the clusters are automatically extracted. |
4 | Data size/Scaling | Limited to small size, indexing and querying is different from clustering. | Any size of data can be clustered. |
5 | Mixed data | Depends upon embeddings, some data cant be used because embeddings wont be generated or useless. | Works for pure text, long form and mixed data |
6 | Compute | High memory and processor power needed. | Works on Low memory and compute. |
7 | Hierarchical Clustering | No | Yes |
8 | Speed | Depends up on the datasize | Fastest on any datasize |
Comments